Pith. sign in

REVIEW 5 cited by

CounTR: Transformer-based Generalised Visual Counting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.13721 v3 pith:A75754N7 submitted 2022-08-29 cs.CV

classification cs.CV
keywords countingexemplarsgeneralisedmodelnumbervisualarbitrarycategories
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we consider the problem of generalised visual object counting, with the goal of developing a computational model for counting the number of objects from arbitrary semantic categories, using arbitrary number of "exemplars", i.e. zero-shot or few-shot counting. To this end, we make the following four contributions: (1) We introduce a novel transformer-based architecture for generalised visual object counting, termed as Counting Transformer (CounTR), which explicitly capture the similarity between image patches or with given "exemplars" with the attention mechanism;(2) We adopt a two-stage training regime, that first pre-trains the model with self-supervised learning, and followed by supervised fine-tuning;(3) We propose a simple, scalable pipeline for synthesizing training images with a large number of instances or that from different semantic categories, explicitly forcing the model to make use of the given "exemplars";(4) We conduct thorough ablation studies on the large-scale counting benchmark, e.g. FSC-147, and demonstrate state-of-the-art performance on both zero and few-shot settings.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Depth-Guided Video Object Counting in Crowded Scenes

    cs.CV 2026-08 conditional novelty 7.0 of 10

    A RGB-D video counting method with depth-based feature fusion and occlusion-adaptive tracking reduces counting errors in crowded scenes, validated on a new dataset.

  2. Spatially-Aware Class-Agnostic Object Counting

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A reference-free class-agnostic counter that combines multi-layer ViT features, DPT reassembly, and FeatUp spatial refinement achieves 12.39 MAE on FSC-147 and 6.27 MAE on CARPK.

  3. CountZES: Counting via Zero-Shot Exemplar Selection

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A training-free, three-stage exemplar-selection pipeline improves zero-shot object counting across natural, aerial, and medical images when compared with other inference-only methods.

  4. Text-promptable Object Counting via Quantity Awareness Enhancement

    cs.CV 2025-07 conditional novelty 6.0 of 10

    QUANet improves text-promptable object counting by training with quantity-specified text prompts, a vision-text quantity alignment loss, and a dual-stream density decoder.

  5. SFOOD: A Multimodal Benchmark for Comprehensive Food Attribute Analysis Beyond RGB with Spectral Insights

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SFOOD combines existing food datasets with self-collected hyperspectral images to create a six-task benchmark, and its evaluations suggest spectral bands improve sweetness and herbal classification while current model...

Pith tools