Pith. sign in

REVIEW 2 cited by

Zero-Shot Object Counting with Language-Vision Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.13097 v1 pith:LSV2EWQX submitted 2023-09-22 cs.CV

classification cs.CV
keywords countingobjectclassexemplarsproposeclass-agnosticcontainingenables
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Class-agnostic object counting aims to count object instances of an arbitrary class at test time. It is challenging but also enables many potential applications. Current methods require human-annotated exemplars as inputs which are often unavailable for novel categories, especially for autonomous systems. Thus, we propose zero-shot object counting (ZSC), a new setting where only the class name is available during test time. This obviates the need for human annotators and enables automated operation. To perform ZSC, we propose finding a few object crops from the input image and use them as counting exemplars. The goal is to identify patches containing the objects of interest while also being visually representative for all instances in the image. To do this, we first construct class prototypes using large language-vision models, including CLIP and Stable Diffusion, to select the patches containing the target objects. Furthermore, we propose a ranking model that estimates the counting error of each patch to select the most suitable exemplars for counting. Experimental results on a recent class-agnostic counting dataset, FSC-147, validate the effectiveness of our method.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Contrastive Learning for Referring Expression Counting

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A contrastive training loss that groups the image regions most similar to a referring expression, using the true object count, improves counting accuracy by over 22% on REC-8K.

  2. Expanding Zero-Shot Object Counting with Rich Prompts

    cs.CV 2025-05 conditional novelty 5.0 of 10

    RichCount improves zero-shot object counting by enriching text prompts with MLLM-generated descriptions and aligning them to CLIP visual features, achieving state-of-the-art mean absolute error on three counting benchmarks.

Pith tools