Pith. sign in

REVIEW 4 cited by

ChatGPT Asks, BLIP-2 Answers: Automatic Questioning Towards Enriched Visual Descriptions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.06594 v1 pith:DHPQCHGL submitted 2023-03-12 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords chatcaptionerblip-2imagequestionschatgptquestioningacquiringanswers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Asking insightful questions is crucial for acquiring knowledge and expanding our understanding of the world. However, the importance of questioning has been largely overlooked in AI research, where models have been primarily developed to answer questions. With the recent advancements of large language models (LLMs) like ChatGPT, we discover their capability to ask high-quality questions when provided with a suitable prompt. This discovery presents a new opportunity to develop an automatic questioning system. In this paper, we introduce ChatCaptioner, a novel automatic-questioning method deployed in image captioning. Here, ChatGPT is prompted to ask a series of informative questions about images to BLIP-2, a strong vision question-answering model. By keeping acquiring new visual information from BLIP-2's answers, ChatCaptioner is able to generate more enriched image descriptions. We conduct human-subject evaluations on common image caption datasets such as COCO, Conceptual Caption, and WikiArt, and compare ChatCaptioner with BLIP-2 as well as ground truth. Our results demonstrate that ChatCaptioner's captions are significantly more informative, receiving three times as many votes from human evaluators for providing the most image information. Besides, ChatCaptioner identifies 53% more objects within the image than BLIP-2 alone measured by WordNet synset matching. Code is available at https://github.com/Vision-CAIR/ChatCaptioner

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning

    cs.RO 2025-10 conditional novelty 6.0 of 10

    A four-camera VLA navigation model trained by distilling multiple RL experts achieves strong simulation performance and qualitative real-world transfer.

  2. ReME: A Data-Centric Framework for Training-Free Open-Vocabulary Segmentation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    ReME builds a cleaned, synonym-enriched segment-text reference set from real images and shows that simple similarity retrieval on it beats 14 prior training-free open-vocabulary segmentation methods across ten benchmarks.

  3. Adapting Lightweight Vision Language Models for Radiological Visual Question Answering

    cs.CV 2025-06 reject novelty 4.0 of 10

    A 3B PaliGemma model fine-tuned with synthetic QA pairs and two-stage training reaches 41.5% accuracy on open-ended radiology VQA, about 15 points below LLaVA-Med.

  4. Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges

    cs.CV 2025-07 conditional novelty 2.0 of 10

    A structured survey of prompt engineering methods for the Segment Anything Model, covering geometric, textual, and multimodal prompts and their applications.

Pith tools