REVIEW 9 cited by
Chain of Thought Prompt Tuning in Vision Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Language-Image Pre-training has demonstrated promising results on zero-shot and few-shot downstream tasks by prompting visual models with natural language prompts. However, most recent studies only use a single prompt for tuning, neglecting the inherent step-to-step cognitive reasoning process that humans conduct in complex task settings, for example, when processing images from unfamiliar domains. Chain of Thought is a simple and effective approximation to human reasoning process and has been proven useful for natural language processing (NLP) tasks. Based on this cognitive intuition, we believe that conducting effective reasoning is also an important problem in visual tasks, and a chain of thought could be a solution to this problem. In this work, we propose a novel chain of thought prompt tuning for vision-language modeling. Extensive experiments show that our method not only generalizes better in image classification tasks, has greater transferability beyond a single dataset, and has stronger domain generalization performance, but also performs much better in imagetext retrieval and visual question answering, which require more reasoning capabilities. We are the first to successfully adapt chain-of-thought prompting that combines visual and textual embeddings. We will release our codes
Forward citations
Cited by 9 Pith papers
-
Can Vision-Language Models Assess Proxemic Risk from Egocentric Robot Images?
Vision-language models classify proxemic danger in egocentric robot images only slightly better than random, except Qwen-VL which detects high-danger scenes with high recall that is not tied to accurate person localization.
-
AVA-VLM: Adaptive Visual Attention-Vision Language Model for In-the-Wild Construction Site Monitoring
AVA-VLM reduces visual-token usage by 69% while improving PPE-violation F1 by 13 points over direct-QA baselines by training a VLM to adaptively crop high-resolution local regions from a downsampled global image.
-
Salience Adjustment for Context-Based Emotion Recognition
A salience-weighted Bayesian cue integration formula improves automated emotion recognition on the Split-Steal corpus, but the weighting is fit to the evaluation set and no held-out validation is provided.
-
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
Argus adds explicit language-guided visual attention to multimodal LLMs by grounding questions to bounding boxes and re-engaging those regions, improving vision-centric reasoning and grounding accuracy.
-
Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
A survey organizes multimodal reasoning research into a staged roadmap and proposes native large multimodal reasoning models that unify perception, generation, and agentic planning.
-
XiHeFusion: Harnessing Large Language Models for Science Communication in Nuclear Fusion
XiHeFusion is a Qwen2.5-14B model fine-tuned on 1.2 million fusion knowledge pairs to answer nuclear fusion questions for science communication.
-
Explainability for Vision Foundation Models: A Survey
A structured review of 122 papers on explainability for vision foundation models, with a taxonomy and the finding that quantitative evaluation is rare (36%).
-
A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future
A historical review that organizes multimodal explainability methods into four chronological eras and three explainability types, extending coverage to generative LLMs.
-
Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey
A survey maps the field of MLLM explainability and interpretability into data, model, and training and inference perspectives.
Discussion (0). Continue with ORCID to comment.