Pith. sign in

REVIEW 4 major objections 7 minor 53 references

Seeing the Undefined: Chain-of-Action for Generative Semantic Labels

T0 review · 4 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read The paper claims that Chain-of-Action (CoA), a five-step zero-shot prompting pipeline, makes vision-language models generate broader and more accurate semantic labels without any predefined label set.

desk verdict CoA is a plausible prompting pipeline for a useful new task, but the two proposed metrics both reward label quantity over label quality, so the headline accuracy gains are not yet established. read the letter →

arxiv 2411.17406 v2 pith:CCDFJS2U submitted 2024-11-26 cs.CV

classification cs.CV
keywords GenerativeSemanticLabelsChain-of-ActionVision-LanguageModelZero-shotpromptingOpen-vocabularyclassificationMulti-labelimagerecognitionVocabulary-unknownlabelspacesComposite
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper defines a new task, Generative Semantic Labels (GSLs): given an image, produce a comprehensive set of semantic labels—objects, scenes, attributes, relationships—without any predefined label vocabulary. It then proposes Chain-of-Action (CoA), a zero-shot prompting pipeline that breaks label generation into five sequential steps: caption the image, self-correct the entity list, describe appearances, infer relationships, and synthesize final labels. The paper claims CoA consistently outperforms VQA-based and caption-based baselines built on BLIP-2, InstructBLIP, LLaVA, and MiniGPT-4 across VOC, COCO, NUS, and a collected open-vocabulary dataset. If true, the result matters because it offers a training-free way to make vision-language models usable in open-ended domains where label sets are unknown and often composite. The paper also proposes two evaluation metrics, semantic accuracy and semantic comprehensiveness, as a reusable protocol for this task.

What carries the argument

The load-bearing mechanism is the Chain-of-Action (CoA) prompt sequence, a five-step zero-shot protocol: (1) Caption Action generates a one-sentence image summary and an initial entity list; (2) Self-Correct Action filters that list with targeted yes/no questions; (3) Appearance Action extracts per-entity attributes; (4) Relationship Action infers spatial and interaction links; (5) Final Action integrates the accumulated context into the output labels. Each action hands its enriched context to the next, so the VLM is never asked to infer everything from a single prompt. The paper's ablations treat the five-action chain as the unit under test and attribute the performance gain to the progressive accumulation of context.

What would settle it

Take a random sample of images from VOC and COCO, have independent human annotators judge each predicted label produced by CoA and by the best caption baseline as present or absent in the image, and compare precision and recall. If CoA's margin over the baseline disappears or reverses under human scoring, the reported Macc/Mcom advantages are artifacts of trusting RAM and CLIP as the arbiters of correctness and coverage.

Watch

Extended reading notes

Core claim

The central discovery, stated on the paper's own terms, is that a vision-language model can be guided to 'see the undefined' by decomposing label generation into a chain of progressively enriched actions rather than asking for labels in one shot. CoA first obtains a broad one-sentence caption and extracts an initial object list, then verifies each entity with a yes/no self-correction query, then gathers appearance details and inter-entity relationships, and finally merges all of it into the label set. The paper reports that this full chain exceeds the second-best method by +11.56% in semantic accuracy and +6.01% in semantic comprehensiveness on VOC, with similar gains on COCO, NUS, and a real-world open-vocabulary dataset, and that the gains hold when the base model is changed to LLaVA-v1.6-7B. The paper also shows that multiple interactions beat a single merged interaction, and that caption-first strategies beat direct VQA queries.

Load-bearing premise

The central claim rests on trusting RAM's automatic tagger and CLIP's text-image similarity as faithful proxies for human judgments of label correctness and completeness.

Editorial extensions

If this is right

  • Zero-shot prompting alone can substantially improve multi-label generation across multiple VLM families, making the approach usable without training or external label databases.
  • Decomposing an open-ended generation task into caption, verification, attribute, and relationship sub-questions is a transferable recipe for other vision-language tasks.
  • Caption-based grounding outperforms direct VQA queries for label generation, suggesting broad scene description should precede fine-grained interrogation.
  • The proposed Macc and Mcom metrics give future GSLs work a shared evaluation protocol, even though the protocol relies on automatic scorers.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: The yes/no self-correction step probably does heavy lifting against caption hallucination; the paper does not isolate it, but a per-action hallucination audit would settle it.
  • Inference: The Mcom metric may reward longer label lists if CLIP similarity rises with prompt length, so a length-controlled variant (same number of labels per comparison) would be a sturdier evaluation.
  • Inference: CoA's context chaining could transfer to video captioning or embodied agents, where previous-frame descriptions play the role of the caption step; the paper does not test those settings.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper introduces Generative Semantic Labels (GSLs), a task requiring a vision-language model to produce an open-ended set of composite semantic labels for an image, and proposes Chain-of-Action (CoA), a zero-shot five-step prompting pipeline (caption, self-correction, appearance, relationship, final integration) that progressively enriches context. The authors evaluate CoA on VOC, COCO, NUS-WIDE, and a self-collected OVD dataset, using two proposed metrics: Mcom (CLIP-based comparison of predicted vs. human prompt similarity) and Macc (RAM-confidence sum over predicted labels). They report consistent gains over VQA and captioning baselines across several VLMs, with headline improvements of +11.56% in Macc and +6.01% in Mcom on VOC, and an ablation table that attributes the gains to each action.

Significance. If the evaluation were reliable, the paper would make a useful contribution: GSLs is a sensible task formalization, the prompting strategy is simple and model-agnostic, the code is released, and improvements are reported across multiple base models and datasets. The idea of chaining caption, verification, appearance, and relationship actions is a reasonable way to elicit more structured outputs from VLMs. However, the two proposed metrics are not validated as measures of label correctness or comprehensiveness, and the numeric claims contain inconsistencies with the tables, so the central contribution is currently not established.

major comments (4)
  1. [4.1] The Macc metric is internally inconsistent: the prose states that 'one point is added for each correct prediction, and one point is subtracted for each incorrect prediction,' but the displayed formula AS_i = Σ_j Sigmoid(RAM(Y_j^i)) is a sum of positive sigmoid scores with no penalty term. As written, every emitted label adds a positive quantity, so Macc increases with the number of predicted labels regardless of their correctness; CoA is exactly the configuration that emits the most labels. The headline +11.56% Macc advantage (Section 4.2) may therefore reflect a length effect, and the central claim of improved label accuracy is not established. Please provide a corrected formula that implements the stated scoring rule, or validate the sigmoid sum against human judgments.
  2. [4.2] The reported headline improvements are not consistent with the tables. For VOC, the text claims '+11.56% in Macc and +6.01% in Mcom' over the second-best method, but Table 1(a) shows the best baseline Macc is 75.34 (Caption LLaVA) and the best baseline Mcom is 70.89 (Caption InstructBLIP); the corresponding differences are 2.15 and 11.56 respectively. For COCO, the text reports '+10.77% in Macc and +6.21% in Mcom', but Table 1(b) yields +10.77% in Mcom and +6.21% in Macc. For NUS, the text reports '+13.4% in Macc and +4.45% in Mcom', while Table 2 yields +12.05% in Mcom and +4.45% in Macc. Please re-check the calculations and state precisely which baseline is being used.
  3. [4.4 (Table 3)] The ablation table does not support the claim that each added action contributes improvements. In Table 3, row (iv) (Actions 1+2+3+5) decreases Vehicles Mcom from 80.14 in row (iii) to 79.63, and row (v) (full CoA) decreases Vehicles Macc from 75.84 in row (iv) to 75.18; these negative changes are not reported in the parenthetical deltas. The statement in Section 4.4 that 'each added component contributes to significant improvements' is therefore not supported by the displayed numbers. Please report all deltas and provide a statistical or consistency analysis.
  4. [4.3 (Table 5)] The results for the real-world open-vocabulary dataset appear to be identical to the NUS-WIDE Nature subset in Table 2 (e.g., VQA BLIP-2 is 70.76/60.54, VQA InstructBLIP is 64.40/62.70, Caption LLaVA is 74.76/61.71). Moreover, the claimed gains do not match the table: from Table 5, the best baseline Mcom is 77.84 (Caption MiniGPT-4), giving a CoA advantage of 5.46 rather than 6.7%, and the best baseline Macc is 74.76 (Caption LLaVA), giving an advantage of 8.54 rather than 7.73%. Please clarify the provenance of the OVD numbers and correct the claims.
minor comments (7)
  1. [4.1] In the formula for AS_i, 'Sigmod' should be 'Sigmoid'.
  2. [4.1] The text says Mcom 'assess[es] the coverage of predicted labels,' but the metric is a relative comparison of CLIP similarities between the predicted prompt and the manual-annotation prompt; this should be described as a relative preference, not an absolute coverage score.
  3. [4.2] The term 'second-best method' is ambiguous; specify whether it means the best baseline (excluding CoA) or the second-best of all methods, and use that definition consistently throughout.
  4. [4.4] The header of Table 3 contains a typo: 'datatset' should be 'dataset'.
  5. [3.3] The 'filtering strategy' that extracts the initial object list from the caption is not described; provide the extraction algorithm or examples for reproducibility.
  6. [References] References [49] and [50] are the same paper (MiniGPT-4) and should be merged.
  7. [5] The limitation statement discusses coarse- and fine-grained labels but does not mention the reliance on RAM and CLIP as proxy ground truth; add a discussion of the limitations of the proposed evaluation metrics.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: CoA is a zero-shot prompting pipeline; the Macc/Mcom metric concerns are validity issues, not circularity.

full rationale

The CoA method is a zero-shot prompting sequence (Caption, Self-Correct, Appearance, Relationship, Final) applied to a frozen VLM. No fitted parameter is renamed as a prediction: the only tuning is the choice of the Caption instruction from a predefined set using a small auxiliary dataset (§3.3), which is a prompt-selection step, not a fitted quantity used as evidence. There are no load-bearing self-citations; RAM [43] and CLIP [26] are external models, not the authors' prior work, and are not invoked to justify the method's design. The concern that 'AS_i = Σ_j Sigmoid(RAM(Y_j^i))' in §4.1 is inconsistent with the prose scoring rule, and the note in the Fig. 3 caption that 'dataset-provided labels are used as human annotations, which are relatively limited' for Mcom, are legitimate threats to the validity of the reported numbers; however, they concern whether the metrics measure what they claim, not whether a derivation reduces to its inputs. The label set is generated from the image by prompting, and the reported improvements, even if inflated by metric bias, are not equivalent by construction to the method's inputs. Therefore the paper's claimed derivation chain is self-contained with respect to circularity; any weaknesses should be addressed as evaluation validity or correctness risks, not circularity.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The paper introduces no invented physical entities. Its claims depend on two black-box evaluators, CLIP and RAM, being valid judges of label quality, on the VLM's internal knowledge being a sufficient label space, and on undisclosed choices in prompt selection and filtering. The free parameters are discrete choices and thresholds that could change the reported numbers.

free parameters (3)
  • RAM confidence threshold for Macc = unspecified
    The Macc metric filters predicted labels by RAM confidence thresholds, but no threshold value is reported; the accuracy numbers depend on this cut-off.
  • Caption prompt selection = optimal prompt from a predefined set, chosen on an auxiliary dataset
    Section 3.3(i) states the caption instruction is selected by optimization over a small auxiliary dataset, a discrete free choice affecting every later step.
  • Object list filtering strategy = unspecified
    The initial object list is extracted from the caption with a 'filtering strategy' that is never specified; different parsers or rules would change the final label set.
assumptions (5)
  • domain assumption CLIP cosine similarity to the image is a valid proxy for semantic comprehensiveness of a label set.
    Mcom in Section 4.1 scores predicted labels by whether CLIP prefers 'This image contains [predicted]' over 'This image contains [human labels]'; this assumes CLIP ranks label sets by relevance.
  • domain assumption RAM confidence scores are a valid measure of individual label correctness.
    Macc in Section 4.1 sums RAM's sigmoid confidence scores to define accuracy, assuming RAM's confidence corresponds to true correctness without validation on human labels.
  • domain assumption The VLM's internal parametric knowledge constitutes a sufficiently complete semantic space S.
    Section 3.2 states S is 'inherently the knowledge base of the VLMs'; the method relies on this rather than an external label database.
  • domain assumption The auxiliary dataset used for prompt selection is unbiased and does not overlap with test sets.
    Section 3.3(i) describes optimization on a small auxiliary dataset but does not identify it; overlap could leak test information into prompt choice.
  • ad hoc to paper Human-annotated labels are 'relatively limited' yet suitable as the comparison baseline for Mcom.
    Section 4.1 and Figure 3 caption state dataset labels are used as human annotations but are relatively limited; Mcom still treats them as the reference to beat.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Seeing the Undefined: Chain-of-Action for Generative Semantic Labels." pith.science (2026). https://pith.science/paper/CCDFJS2U

@misc{pith2026241117406,
  author       = {Pith},
  title        = {Pith review of: Seeing the Undefined: Chain-of-Action for Generative Semantic Labels},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CCDFJS2U}},
  note         = {Machine review of arXiv:2411.17406}
}
read the original abstract

Recent advances in vision-language models (VLMs) have demonstrated remarkable capabilities in image classification by leveraging predefined sets of labels to construct text prompts for zero-shot reasoning. However, these approaches face significant limitations in undefined domains, where the label space is vocabulary-unknown and composite. We thus introduce Generative Semantic Labels (GSLs), a novel task that aims to predict a comprehensive set of semantic labels for an image without being constrained by a predefined labels set. Unlike traditional zero-shot classification, GSLs generates multiple semantic-level labels, encompassing objects, scenes, attributes, and relationships, thereby providing a richer and more accurate representation of image content. In this paper, we propose Chain-of-Action (CoA), an innovative method designed to tackle the GSLs task. CoA is motivated by the observation that enriched contextual information significantly improves generative performance during inference. Specifically, CoA decomposes the GSLs task into a sequence of detailed actions. Each action extracts and merges key information from the previous step, passing enriched context to the next, ultimately guiding the VLM to generate comprehensive and accurate semantic labels. We evaluate the effectiveness of CoA through extensive experiments on widely-used benchmark datasets. The results demonstrate significant improvements across key performance metrics, validating the capability of CoA to generate accurate and contextually rich semantic labels. Our work not only advances the state-of-the-art in generative semantic labels but also opens new avenues for applying VLMs in open-ended and dynamic real-world scenarios.

Figures

Figures reproduced from arXiv: 2411.17406 by the authors.

Figure 1
Figure 1. An example of the Generative Semantic Labels [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Pipeline of CoA. First, we utilize Caption Action to generate a comprehensive image description and extract an initial object list. Self-Correct Action refines this list, followed by Appearance Action and Relationship Action to capture detailed appearance features and relationships. Finally, we integrate these responses to create enriched context, guiding VLMs to achieve final semantic labels. in GSLs task. To addre… view at source ↗
Figure 3
Figure 3. Comparison of various settings, including VQA [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: The instruction templates defined within the vari [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Comparison of multiple interaction (CoA) with single merge interaction on the VOC datatset. [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

53 extracted references · 37 canonical work pages

  1. [1]

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al

  2. [2]

    Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015. Vqa: Visual question answering. In Proceedings of the IEEE International Conference on Computer Vision. 2425–2433

  3. [3]

    Amir Bar, Yossi Gandelsman, Trevor Darrell, Amir Globerson, and Alexei A. Efros

  4. [4]

    Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, et al. 2024. Graph of thoughts: Solving elaborate problems with large language models. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 17682–17690

  5. [5]

    InProceedings of Annual Conference on Neural Information Processing Systems

    Visual Prompting via Image Inpainting. InProceedings of Annual Conference on Neural Information Processing Systems

  6. [6]

    Tat-Seng Chua, Jinhui Tang, Richang Hong, Haojie Li, Zhiping Luo, and Yantao Zheng. 2009. Nus-wide: a real-world web image database from national university of singapore. InProceedings of the ACM international conference on image and video retrieval. 1–9

  7. [7]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin...

  8. [8]

    Wenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale Fung, and Steven C. H. Hoi. 2023. In- structBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning. InProceedings of Advances in Neural Information Processing Systems

Show all 53 references
  1. [9]

    Alessandro Conti, Enrico Fini, Massimiliano Mancini, Paolo Rota, Yiming Wang, and Elisa Ricci. 2023. Vocabulary-free Image Classification. InProceedings of Annual Conference on Neural Information Processing Systems

  2. [10]

    Zhengxiao Du, Yujie Qian, Xiao Liu, Ming Ding, Jiezhong Qiu, Zhilin Yang, and Jie Tang. 2022. GLM: General Language Model Pretraining with Autoregressive Blank Infilling. InProceedings of Annual Meeting of the Association for Computational Linguistics. 320–335

  3. [11]

    Danny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch, Aakanksha Chowdh- ery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, Wenlong Huang, Yevgen Chebotar, Pierre Sermanet, Daniel Duckworth, Sergey Levine, Vincent Vanhoucke, Karol Hausman, Marc Toussai...

  4. [12]

    Xiang Gao and Kamalika Das. 2024. Customizing Language Model Responses with Contrastive In-Context Learning. InProceedings of the AAAI Conference on Artificial Intelligence. 18039–18046

  5. [13]

    M Everingham, L Van Gool, CKI Williams, J Winn, and A Zisserman. 2012. The PASCAL visual object classes challenge 2012 (VOC2012) results. 2012 http://www. pascal-network. org/challenges. InVOC/voc2012/workshop/index. html

  6. [14]

    Drew A Hudson and Christopher D Manning. 2019. Gqa: A new dataset for real- world visual reasoning and compositional question answering. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 6700–6709

  7. [15]

    Shashank Goel, Hritik Bansal, Sumit Bhatia, Ryan Rossi, Vishwa Vinay, and Aditya Grover. 2022. Cyclip: Cyclic contrastive language-image pretraining. Advances in Neural Information Processing Systems35 (2022), 6704–6719

  8. [16]

    Bin Lei, Chunhua Liao, Caiwen Ding, et al. 2023. Boosting logical reasoning in large language models through a new framework: The graph of thought.arXiv preprint arXiv:2308.08614(2023)

  9. [17]

    Ryuto Koike, Masahiro Kaneko, and Naoaki Okazaki. [n. d.]. OUTFOX: LLM- Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples. InProceedings of the AAAI Conference on Artificial Intelli- gence. 21258–21266

  10. [18]

    Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. InProceedings of International Conference on Machine Learning. PMLR, 12888–12900

  11. [19]

    Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. 2023. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. InProceedings of International Conference on Machine Learning. PMLR, 19730–19742

  12. [20]

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Improved baselines with visual instruction tuning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 26296–26306

  13. [21]

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. InComputer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Procee...

  14. [22]

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruc- tion tuning.Advances in Neural Information Processing Systems36 (2024)

  15. [23]

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. 2024. Improved Baselines with Visual Instruction Tuning. InIEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2024, Seattle, W A, USA, June 16-22, 2024. 26286–26296. https://doi.org/10.1109/CVPR52733.2024.02484

  16. [24]

    Chancharik Mitra, Brandon Huang, Trevor Darrell, and Roei Herzig. 2024. Com- positional Chain-of-Thought Prompting for Large Multimodal Models. InPro- ceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition. 14420–14431

  17. [25]

    Haotian Liu, Kilho Son, Jianwei Yang, Ce Liu, Jianfeng Gao, Yong Jae Lee, and Chunyuan Li. 2023. Learning Customized Visual Models with Retrieval- Augmented Knowledge. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 15148–15158

  18. [26]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. InProceedings of International Conference on ...

  19. [27]

    Zhijie Nie, Richong Zhang, Zhongyuan Wang, and Xudong Liu. 2024. Code-Style In-Context Learning for Knowledge-Based Question Answering. InProceedings of the AAAI Conference on Artificial Intelligence. 18833–18841

  20. [28]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research21, 140 (2020), 1–67

  21. [29]

    Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners.OpenAI blog 1, 8 (2019), 9

  22. [30]

    Dianmo Sheng, Dongdong Chen, Zhentao Tan, Qiankun Liu, Qi Chu, Jianmin Bao, Tao Gong, Bin Liu, Shengwei Xu, and Nenghai Yu. 2024. Towards More Unified In-Context Visual Understanding. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 13362–13372

  23. [31]

    Tanik Saikh, Tirthankar Ghosal, Amish Mittal, Asif Ekbal, and Pushpak Bhat- tacharyya. 2022. Scienceqa: A novel resource for question answering on scholarly articles.International Journal on Digital Libraries23, 3 (2022), 289–301

  24. [32]

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al. 2023. Llama: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971(2023)

  25. [33]

    Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler

    Yi Tay, Mostafa Dehghani, Vinh Q. Tran, Xavier Garcia, Jason Wei, Xuezhi Wang, Hyung Won Chung, Dara Bahri, Tal Schuster, Huaixiu Steven Zheng, Denny Zhou, Neil Houlsby, and Donald Metzler. 2023. UL2: Unifying Language Learning Paradigms. InProceedings of International Confere...

  26. [34]

    Le, Ed H

    Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-Consistency Improves Chain of Thought Reasoning in Language Models. InProceedings of International Conference on Learning Representations

  27. [35]

    Lei Wang, Yi Hu, Jiabang He, Xing Xu, Ning Liu, Hui Liu, and Heng Tao Shen

  28. [36]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits reasoning in large language models.Advances in Neural Information Processing Systems35 (2022), 24824–24837

  29. [37]

    Junyi Yao, Yijiang Liu, Zhen Dong, Mingfei Guo, Helan Hu, Kurt Keutzer, Li Du, Daquan Zhou, and Shanghang Zhang. 2024. PromptCoT: Align Prompt Distribu- tion via Adapted Chain-of-Thought. InProceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition. 7027–7037

  30. [38]

    Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M

    Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022. Finetuned Language Models are Zero-Shot Learners. InProceedings of International Conference on Learning Representations

  31. [39]

    Yao Yao, Zuchao Li, and Hai Zhao. 2023. Beyond Chain-of-Thought, Ef- fective Graph-of-Thought Reasoning in Language Models.arXiv preprint arXiv:2305.16582(2023)

  32. [40]

    Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, Zhiyuan Liu, Peng Zhang, Yuxiao Dong, and Jie Tang. 2023. GLM-130B: An Open Bilingual Pre-trained...

  33. [41]

    Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Tom Griffiths, Yuan Cao, and Karthik Narasimhan. 2024. Tree of thoughts: Deliberate problem solving with large language models.Advances in Neural Information Processing Systems36 MM ’25, October 27–31, 2025, Dublin, Ireland Men...

  34. [42]

    Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022. Opt: Open pre-trained transformer language models.arXiv preprint arXiv:2205.01068 (2022)

  35. [43]

    Youcai Zhang, Xinyu Huang, Jinyu Ma, Zhaoyang Li, Zhaochuan Luo, Yanchun Xie, Yuzhuo Qin, Tong Luo, Yaqian Li, Shilong Liu, Yandong Guo, and Lei Zhang

  36. [44]

    Jingyi Zhang, Jiaxing Huang, Sheng Jin, and Shijian Lu. 2024. Vision-Language Models for Vision Tasks: A Survey.IEEE Trans. Pattern Anal. Mach. Intell.46, 8 (2024), 5625–5644

  37. [45]

    Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. 2023. Automatic Chain of Thought Prompting in Large Language Models. InProceedings of International Conference on Learning Representations

  38. [46]

    Zhuosheng Zhang, Aston Zhang, Mu Li, Hai Zhao, George Karypis, and Alex Smola. 2024. Multimodal Chain-of-Thought Reasoning in Language Models. Trans. Mach. Learn. Res.2024 (2024)

  39. [47]

    InIEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, CVPR 2024 - Workshops, Seattle, W A, USA, June 17-18, 2024

    Recognize Anything: A Strong Image Tagging Model. InIEEE/CVF Confer- ence on Computer Vision and Pattern Recognition, CVPR 2024 - Workshops, Seattle, W A, USA, June 17-18, 2024. 1724–1732

  40. [48]

    Yuanhan Zhang, Kaiyang Zhou, and Ziwei Liu. 2023. What Makes Good Examples for Visual In-Context Learning?. InProceedings of Annual Conference on Neural Information Processing Systems

  41. [50]

    Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, and Mohamed Elhoseiny. 2024. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models. InProceedings of International Conference on Learning Repre- sentations

  42. [51]

    Ge Zheng, Bin Yang, Jiajin Tang, Hong-Yu Zhou, and Sibei Yang. 2023. Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neural Information Processing Systems36 (2023), 5168–5191

  43. [52]

    Le, and Ed H

    Denny Zhou, Nathanael Schärli, Le Hou, Jason Wei, Nathan Scales, Xuezhi Wang, Dale Schuurmans, Claire Cui, Olivier Bousquet, Quoc V. Le, and Ed H. Chi. 2023. Least-to-Most Prompting Enables Complex Reasoning in Large Language Models. InProceedings of International Conference o...

  44. [2022]

    Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems35 (2022), 23716–23736

  45. [2024]

    InProceedings of the AAAI Conference on Artificial Intelligence, Vol

    T-sciq: Teaching multimodal chain-of-thought reasoning via large lan- guage model signals for science question answering. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 19162–19170

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.