REVIEW 4 major objections 7 minor 53 references
Seeing the Undefined: Chain-of-Action for Generative Semantic Labels
T0 review · 4 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper claims that Chain-of-Action (CoA), a five-step zero-shot prompting pipeline, makes vision-language models generate broader and more accurate semantic labels without any predefined label set.
desk verdict CoA is a plausible prompting pipeline for a useful new task, but the two proposed metrics both reward label quantity over label quality, so the headline accuracy gains are not yet established. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the Chain-of-Action (CoA) prompt sequence, a five-step zero-shot protocol: (1) Caption Action generates a one-sentence image summary and an initial entity list; (2) Self-Correct Action filters that list with targeted yes/no questions; (3) Appearance Action extracts per-entity attributes; (4) Relationship Action infers spatial and interaction links; (5) Final Action integrates the accumulated context into the output labels. Each action hands its enriched context to the next, so the VLM is never asked to infer everything from a single prompt. The paper's ablations treat the five-action chain as the unit under test and attribute the performance gain to the progressive accumulation of context.
What would settle it
Take a random sample of images from VOC and COCO, have independent human annotators judge each predicted label produced by CoA and by the best caption baseline as present or absent in the image, and compare precision and recall. If CoA's margin over the baseline disappears or reverses under human scoring, the reported Macc/Mcom advantages are artifacts of trusting RAM and CLIP as the arbiters of correctness and coverage.
Extended reading notes
Core claim
The central discovery, stated on the paper's own terms, is that a vision-language model can be guided to 'see the undefined' by decomposing label generation into a chain of progressively enriched actions rather than asking for labels in one shot. CoA first obtains a broad one-sentence caption and extracts an initial object list, then verifies each entity with a yes/no self-correction query, then gathers appearance details and inter-entity relationships, and finally merges all of it into the label set. The paper reports that this full chain exceeds the second-best method by +11.56% in semantic accuracy and +6.01% in semantic comprehensiveness on VOC, with similar gains on COCO, NUS, and a real-world open-vocabulary dataset, and that the gains hold when the base model is changed to LLaVA-v1.6-7B. The paper also shows that multiple interactions beat a single merged interaction, and that caption-first strategies beat direct VQA queries.
Load-bearing premise
The central claim rests on trusting RAM's automatic tagger and CLIP's text-image similarity as faithful proxies for human judgments of label correctness and completeness.
Editorial extensions
If this is right
- Zero-shot prompting alone can substantially improve multi-label generation across multiple VLM families, making the approach usable without training or external label databases.
- Decomposing an open-ended generation task into caption, verification, attribute, and relationship sub-questions is a transferable recipe for other vision-language tasks.
- Caption-based grounding outperforms direct VQA queries for label generation, suggesting broad scene description should precede fine-grained interrogation.
- The proposed Macc and Mcom metrics give future GSLs work a shared evaluation protocol, even though the protocol relies on automatic scorers.
Reading between the lines
- Inference: The yes/no self-correction step probably does heavy lifting against caption hallucination; the paper does not isolate it, but a per-action hallucination audit would settle it.
- Inference: The Mcom metric may reward longer label lists if CLIP similarity rises with prompt length, so a length-controlled variant (same number of labels per comparison) would be a sturdier evaluation.
- Inference: CoA's context chaining could transfer to video captioning or embodied agents, where previous-frame descriptions play the role of the caption step; the paper does not test those settings.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Generative Semantic Labels (GSLs), a task requiring a vision-language model to produce an open-ended set of composite semantic labels for an image, and proposes Chain-of-Action (CoA), a zero-shot five-step prompting pipeline (caption, self-correction, appearance, relationship, final integration) that progressively enriches context. The authors evaluate CoA on VOC, COCO, NUS-WIDE, and a self-collected OVD dataset, using two proposed metrics: Mcom (CLIP-based comparison of predicted vs. human prompt similarity) and Macc (RAM-confidence sum over predicted labels). They report consistent gains over VQA and captioning baselines across several VLMs, with headline improvements of +11.56% in Macc and +6.01% in Mcom on VOC, and an ablation table that attributes the gains to each action.
Significance. If the evaluation were reliable, the paper would make a useful contribution: GSLs is a sensible task formalization, the prompting strategy is simple and model-agnostic, the code is released, and improvements are reported across multiple base models and datasets. The idea of chaining caption, verification, appearance, and relationship actions is a reasonable way to elicit more structured outputs from VLMs. However, the two proposed metrics are not validated as measures of label correctness or comprehensiveness, and the numeric claims contain inconsistencies with the tables, so the central contribution is currently not established.
major comments (4)
- [4.1] The Macc metric is internally inconsistent: the prose states that 'one point is added for each correct prediction, and one point is subtracted for each incorrect prediction,' but the displayed formula AS_i = Σ_j Sigmoid(RAM(Y_j^i)) is a sum of positive sigmoid scores with no penalty term. As written, every emitted label adds a positive quantity, so Macc increases with the number of predicted labels regardless of their correctness; CoA is exactly the configuration that emits the most labels. The headline +11.56% Macc advantage (Section 4.2) may therefore reflect a length effect, and the central claim of improved label accuracy is not established. Please provide a corrected formula that implements the stated scoring rule, or validate the sigmoid sum against human judgments.
- [4.2] The reported headline improvements are not consistent with the tables. For VOC, the text claims '+11.56% in Macc and +6.01% in Mcom' over the second-best method, but Table 1(a) shows the best baseline Macc is 75.34 (Caption LLaVA) and the best baseline Mcom is 70.89 (Caption InstructBLIP); the corresponding differences are 2.15 and 11.56 respectively. For COCO, the text reports '+10.77% in Macc and +6.21% in Mcom', but Table 1(b) yields +10.77% in Mcom and +6.21% in Macc. For NUS, the text reports '+13.4% in Macc and +4.45% in Mcom', while Table 2 yields +12.05% in Mcom and +4.45% in Macc. Please re-check the calculations and state precisely which baseline is being used.
- [4.4 (Table 3)] The ablation table does not support the claim that each added action contributes improvements. In Table 3, row (iv) (Actions 1+2+3+5) decreases Vehicles Mcom from 80.14 in row (iii) to 79.63, and row (v) (full CoA) decreases Vehicles Macc from 75.84 in row (iv) to 75.18; these negative changes are not reported in the parenthetical deltas. The statement in Section 4.4 that 'each added component contributes to significant improvements' is therefore not supported by the displayed numbers. Please report all deltas and provide a statistical or consistency analysis.
- [4.3 (Table 5)] The results for the real-world open-vocabulary dataset appear to be identical to the NUS-WIDE Nature subset in Table 2 (e.g., VQA BLIP-2 is 70.76/60.54, VQA InstructBLIP is 64.40/62.70, Caption LLaVA is 74.76/61.71). Moreover, the claimed gains do not match the table: from Table 5, the best baseline Mcom is 77.84 (Caption MiniGPT-4), giving a CoA advantage of 5.46 rather than 6.7%, and the best baseline Macc is 74.76 (Caption LLaVA), giving an advantage of 8.54 rather than 7.73%. Please clarify the provenance of the OVD numbers and correct the claims.
minor comments (7)
- [4.1] In the formula for AS_i, 'Sigmod' should be 'Sigmoid'.
- [4.1] The text says Mcom 'assess[es] the coverage of predicted labels,' but the metric is a relative comparison of CLIP similarities between the predicted prompt and the manual-annotation prompt; this should be described as a relative preference, not an absolute coverage score.
- [4.2] The term 'second-best method' is ambiguous; specify whether it means the best baseline (excluding CoA) or the second-best of all methods, and use that definition consistently throughout.
- [4.4] The header of Table 3 contains a typo: 'datatset' should be 'dataset'.
- [3.3] The 'filtering strategy' that extracts the initial object list from the caption is not described; provide the extraction algorithm or examples for reproducibility.
- [References] References [49] and [50] are the same paper (MiniGPT-4) and should be merged.
- [5] The limitation statement discusses coarse- and fine-grained labels but does not mention the reliance on RAM and CLIP as proxy ground truth; add a discussion of the limitations of the proposed evaluation metrics.
Circularity Check
No significant circularity: CoA is a zero-shot prompting pipeline; the Macc/Mcom metric concerns are validity issues, not circularity.
full rationale
The CoA method is a zero-shot prompting sequence (Caption, Self-Correct, Appearance, Relationship, Final) applied to a frozen VLM. No fitted parameter is renamed as a prediction: the only tuning is the choice of the Caption instruction from a predefined set using a small auxiliary dataset (§3.3), which is a prompt-selection step, not a fitted quantity used as evidence. There are no load-bearing self-citations; RAM [43] and CLIP [26] are external models, not the authors' prior work, and are not invoked to justify the method's design. The concern that 'AS_i = Σ_j Sigmoid(RAM(Y_j^i))' in §4.1 is inconsistent with the prose scoring rule, and the note in the Fig. 3 caption that 'dataset-provided labels are used as human annotations, which are relatively limited' for Mcom, are legitimate threats to the validity of the reported numbers; however, they concern whether the metrics measure what they claim, not whether a derivation reduces to its inputs. The label set is generated from the image by prompting, and the reported improvements, even if inflated by metric bias, are not equivalent by construction to the method's inputs. Therefore the paper's claimed derivation chain is self-contained with respect to circularity; any weaknesses should be addressed as evaluation validity or correctness risks, not circularity.
Assumptions & free parameters
free parameters (3)
- RAM confidence threshold for Macc =
unspecified
- Caption prompt selection =
optimal prompt from a predefined set, chosen on an auxiliary dataset
- Object list filtering strategy =
unspecified
assumptions (5)
- domain assumption CLIP cosine similarity to the image is a valid proxy for semantic comprehensiveness of a label set.
- domain assumption RAM confidence scores are a valid measure of individual label correctness.
- domain assumption The VLM's internal parametric knowledge constitutes a sufficiently complete semantic space S.
- domain assumption The auxiliary dataset used for prompt selection is unbiased and does not overlap with test sets.
- ad hoc to paper Human-annotated labels are 'relatively limited' yet suitable as the comparison baseline for Mcom.
Cite this review
Pith. "Pith review of Seeing the Undefined: Chain-of-Action for Generative Semantic Labels." pith.science (2026). https://pith.science/paper/CCDFJS2U
@misc{pith2026241117406,
author = {Pith},
title = {Pith review of: Seeing the Undefined: Chain-of-Action for Generative Semantic Labels},
year = {2026},
howpublished = {\url{https://pith.science/paper/CCDFJS2U}},
note = {Machine review of arXiv:2411.17406}
}
read the original abstract
Recent advances in vision-language models (VLMs) have demonstrated remarkable capabilities in image classification by leveraging predefined sets of labels to construct text prompts for zero-shot reasoning. However, these approaches face significant limitations in undefined domains, where the label space is vocabulary-unknown and composite. We thus introduce Generative Semantic Labels (GSLs), a novel task that aims to predict a comprehensive set of semantic labels for an image without being constrained by a predefined labels set. Unlike traditional zero-shot classification, GSLs generates multiple semantic-level labels, encompassing objects, scenes, attributes, and relationships, thereby providing a richer and more accurate representation of image content. In this paper, we propose Chain-of-Action (CoA), an innovative method designed to tackle the GSLs task. CoA is motivated by the observation that enriched contextual information significantly improves generative performance during inference. Specifically, CoA decomposes the GSLs task into a sequence of detailed actions. Each action extracts and merges key information from the previous step, passing enriched context to the next, ultimately guiding the VLM to generate comprehensive and accurate semantic labels. We evaluate the effectiveness of CoA through extensive experiments on widely-used benchmark datasets. The results demonstrate significant improvements across key performance metrics, validating the capability of CoA to generate accurate and contextually rich semantic labels. Our work not only advances the state-of-the-art in generative semantic labels but also opens new avenues for applying VLMs in open-ended and dynamic real-world scenarios.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al
-
[2]
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In Proceedings of the IEEE International Conference on Computer Vision. 2425–2433
work page 2015
-
[3]
Amir Bar, Yossi Gandelsman, Trevor Darrell, Amir Globerson, and Alexei A. Efros
-
[4]
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, et al. 2024. Graph of thoughts: Solving elaborate problems with large language models. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 17682–17690
work page 2024
-
[5]
InProceedings of Annual Conference on Neural Information Processing Systems
Visual Prompting via Image Inpainting. InProceedings of Annual Conference on Neural Information Processing Systems
-
[6]
Tat-Seng Chua, Jinhui Tang, Richang Hong, Haojie Li, Zhiping Luo, and Yantao Zheng. 2009. Nus-wide: a real-world web image database from national university of singapore. InProceedings of the ACM international conference on image and video retrieval. 1–9
2009
-
[7]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin...
work page 2020
-
[8]
Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven C. H. Hoi. 2023. In- structBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. InProceedings of Advances in Neural Information Processing Systems
work page 2023
Show all 53 references
-
[9]
Alessandro Conti, Enrico Fini, Massimiliano Mancini, Paolo Rota, Yiming Wang, and Elisa Ricci. 2023. Vocabulary-free Image Classification. InProceedings of Annual Conference on Neural Information Processing Systems
2023
-
[10]
Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. GLM: General Language Model Pretraining with Autoregressive Blank Infilling. InProceedings of Annual Meeting of the Association for Computational Linguistics. 320–335
2022
-
[11]
Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdh- ery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussai...
2023
-
[12]
Xiang Gao and Kamalika Das. 2024. Customizing Language Model Responses with Contrastive In-Context Learning. InProceedings of the AAAI Conference on Artificial Intelligence. 18039–18046
2024
-
[13]
M Everingham, L Van Gool, CKI Williams, J Winn, and A Zisserman. 2012. The PASCAL visual object classes challenge 2012 (VOC2012) results. 2012 http://www. pascal-network. org/challenges. InVOC/voc2012/workshop/index. html
2012
-
[14]
Drew A Hudson and Christopher D Manning. 2019. Gqa: A new dataset for real- world visual reasoning and compositional question answering. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 6700–6709
2019
-
[15]
Shashank Goel, Hritik Bansal, Sumit Bhatia, Ryan Rossi, Vishwa Vinay, and Aditya Grover. 2022. Cyclip: Cyclic contrastive language-image pretraining. Advances in Neural Information Processing Systems35 (2022), 6704–6719
2022
-
[16]
Bin Lei, Chunhua Liao, Caiwen Ding, et al. 2023. Boosting logical reasoning in large language models through a new framework: The graph of thought.arXiv preprint arXiv:2308.08614(2023)
2023 arXiv
-
[17]
Ryuto Koike, Masahiro Kaneko, and Naoaki Okazaki. [n. d.]. OUTFOX: LLM- Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples. InProceedings of the AAAI Conference on Artificial Intelli- gence. 21258–21266
-
[18]
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. InProceedings of International Conference on Machine Learning. PMLR, 12888–12900
2022
-
[19]
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. InProceedings of International Conference on Machine Learning. PMLR, 19730–19742
2023
-
[20]
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Improved baselines with visual instruction tuning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 26296–26306
2024
-
[21]
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. InComputer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Procee...
2014
-
[22]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruc- tion tuning.Advances in Neural Information Processing Systems36 (2024)
2024
-
[23]
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Improved Baselines with Visual Instruction Tuning. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, W A, USA, June 16-22, 2024. 26286–26296. https://doi.org/10.1109/CVPR52733.2024.02484
2024
-
[24]
Chancharik Mitra, Brandon Huang, Trevor Darrell, and Roei Herzig. 2024. Com- positional Chain-of-Thought Prompting for Large Multimodal Models. InPro- ceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition. 14420–14431
2024
-
[25]
Haotian Liu, Kilho Son, Jianwei Yang, Ce Liu, Jianfeng Gao, Yong Jae Lee, and Chunyuan Li. 2023. Learning Customized Visual Models with Retrieval- Augmented Knowledge. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 15148–15158
2023
-
[26]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. InProceedings of International Conference on ...
2021
-
[27]
Zhijie Nie, Richong Zhang, Zhongyuan Wang, and Xudong Liu. 2024. Code-Style In-Context Learning for Knowledge-Based Question Answering. InProceedings of the AAAI Conference on Artificial Intelligence. 18833–18841
2024
-
[28]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research21, 140 (2020), 1–67
2020
-
[29]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners.OpenAI blog 1, 8 (2019), 9
2019
-
[30]
Dianmo Sheng, Dongdong Chen, Zhentao Tan, Qiankun Liu, Qi Chu, Jianmin Bao, Tao Gong, Bin Liu, Shengwei Xu, and Nenghai Yu. 2024. Towards More Unified In-Context Visual Understanding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13362–13372
2024
-
[31]
Tanik Saikh, Tirthankar Ghosal, Amish Mittal, Asif Ekbal, and Pushpak Bhat- tacharyya. 2022. Scienceqa: A novel resource for question answering on scholarly articles.International Journal on Digital Libraries23, 3 (2022), 289–301
2022
-
[32]
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971(2023)
2023 arXiv
-
[33]
Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler
Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler. 2023. UL2: Unifying Language Learning Paradigms. InProceedings of International Confere...
2023
-
[34]
Le, Ed H
Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. InProceedings of International Conference on Learning Representations
2023
-
[35]
Lei Wang, Yi Hu, Jiabang He, Xing Xu, Ning Liu, Hui Liu, and Heng Tao Shen
-
[36]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models.Advances in Neural Information Processing Systems35 (2022), 24824–24837
2022
-
[37]
Junyi Yao, Yijiang Liu, Zhen Dong, Mingfei Guo, Helan Hu, Kurt Keutzer, Li Du, Daquan Zhou, and Shanghang Zhang. 2024. PromptCoT: Align Prompt Distribu- tion via Adapted Chain-of-Thought. InProceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition. 7027–7037
2024
-
[38]
Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022. Finetuned Language Models are Zero-Shot Learners. InProceedings of International Conference on Learning Representations
2022
-
[39]
Yao Yao, Zuchao Li, and Hai Zhao. 2023. Beyond Chain-of-Thought, Ef- fective Graph-of-Thought Reasoning in Language Models.arXiv preprint arXiv:2305.16582(2023)
2023 arXiv
-
[40]
Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, Zhiyuan Liu, Peng Zhang, Yuxiao Dong, and Jie Tang. 2023. GLM-130B: An Open Bilingual Pre-trained...
2023
-
[41]
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024. Tree of thoughts: Deliberate problem solving with large language models.Advances in Neural Information Processing Systems36 MM ’25, October 27–31, 2025, Dublin, Ireland Men...
2024
-
[42]
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models.arXiv preprint arXiv:2205.01068 (2022)
2022 arXiv
-
[43]
Youcai Zhang, Xinyu Huang, Jinyu Ma, Zhaoyang Li, Zhaochuan Luo, Yanchun Xie, Yuzhuo Qin, Tong Luo, Yaqian Li, Shilong Liu, Yandong Guo, and Lei Zhang
-
[44]
Jingyi Zhang, Jiaxing Huang, Sheng Jin, and Shijian Lu. 2024. Vision-Language Models for Vision Tasks: A Survey.IEEE Trans. Pattern Anal. Mach. Intell.46, 8 (2024), 5625–5644
2024
-
[45]
Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2023. Automatic Chain of Thought Prompting in Large Language Models. InProceedings of International Conference on Learning Representations
2023
-
[46]
Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. 2024. Multimodal Chain-of-Thought Reasoning in Language Models. Trans. Mach. Learn. Res.2024 (2024)
2024
-
[47]
InIEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, CVPR 2024 - Workshops, Seattle, W A, USA, June 17-18, 2024
Recognize Anything: A Strong Image Tagging Model. InIEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, CVPR 2024 - Workshops, Seattle, W A, USA, June 17-18, 2024. 1724–1732
2024
-
[48]
Yuanhan Zhang, Kaiyang Zhou, and Ziwei Liu. 2023. What Makes Good Examples for Visual In-Context Learning?. InProceedings of Annual Conference on Neural Information Processing Systems
2023
-
[50]
Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2024. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. InProceedings of International Conference on Learning Repre- sentations
2024
-
[51]
Ge Zheng, Bin Yang, Jiajin Tang, Hong-Yu Zhou, and Sibei Yang. 2023. Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neural Information Processing Systems36 (2023), 5168–5191
2023
-
[52]
Le, and Ed H
Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V. Le, and Ed H. Chi. 2023. Least-to-Most Prompting Enables Complex Reasoning in Large Language Models. InProceedings of International Conference o...
2023
-
[2022]
Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems35 (2022), 23716–23736
2022
-
[2024]
InProceedings of the AAAI Conference on Artificial Intelligence, Vol
T-sciq: Teaching multimodal chain-of-thought reasoning via large lan- guage model signals for science question answering. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 19162–19170
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.