Pith. sign in

REVIEW 2 major objections 2 minor 105 references

Count Anything

T0 review · 2 major / 2 minor · reviewed 2026-06-28 · grok-4.3

Pith's one-line read Count Anything shows one model can count objects from text queries across scenes, microscopy, and remote sensing.

desk verdict Count Anything brings text guidance and dual counters to object counting with a new benchmark, but CLOC's construction needs verification to support the generalization claims. read the letter →

arxiv 2605.30846 v1 pith:U4XB4HSG submitted 2026-05-29 cs.CV

classification cs.CV
keywords objectcountingtext-guidedmulti-domaingeneralizationpointpredictionCLOCdatasetdual-granularitycounter
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Object counting has remained split into separate models for crowds, cells, crops and other categories, each failing outside its narrow setting. This paper builds CLOC by merging existing public sources into one collection of 220K images, 619 categories and 15M instances spanning six domains. It trains a single model that takes an image plus a natural-language description and returns a set of points marking each target instance. The architecture splits work between a sparse counter for large isolated objects and a dense counter for tiny crowded ones, then fuses the two outputs without learned parameters. A reader would care because the approach removes the need to select or retrain a specialist counter whenever the visual domain or object type changes.

What carries the argument

Dual-granularity instance enumeration that pairs a Region-level Sparse Counter for large sparse targets with a Pixel-level Dense Counter for small dense targets, fused by Complementary Count Fusion.

What would settle it

Accuracy measurements on a held-out visual domain or object category absent from CLOC training that fall below the performance of existing domain-specific counters.

Watch

Extended reading notes

Core claim

Count Anything is a generalist model for text-guided object counting that replaces density maps with discrete instance points and performs dual-granularity enumeration: a Region-level Sparse Counter supplies anchors for large sparse targets while a Pixel-level Dense Counter predicts dense points for small crowded targets, with point-centric supervision and parameter-free Complementary Count Fusion enabling training on the heterogeneous CLOC collection that covers general scenes, remote sensing, histopathology, cellular microscopy, agriculture and microbiology.

Load-bearing premise

Reorganizing public datasets into CLOC produces a balanced benchmark that tests cross-domain generalization without annotation artifacts or domain leakage.

Editorial extensions

If this is right

  • The same model outperforms prior open-world counting methods on accuracy and multi-domain generalization.
  • Point-centric supervision lets the model learn from mixed annotation formats across source datasets.
  • Output points supply both the count and the spatial locations of every detected instance.
  • CLOC becomes a standard testbed for measuring whether counting models truly cross visual domains.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Point-based counting may replace density maps in settings where instance locations are needed for downstream tasks.
  • Natural-language prompts could let non-specialists request counts inside medical or satellite imagery without writing code.
  • Adding text examples for a new domain might suffice for adaptation instead of collecting fresh labeled images.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The manuscript introduces Count Anything, a text-guided object counting model that employs dual-granularity counters—a Region-level Sparse Counter for large sparse objects and a Pixel-level Dense Counter for small crowded ones—combined via parameter-free Complementary Count Fusion. It also presents CLOC, a new benchmark reorganizing existing datasets into six domains (General Scene, Remote Sensing, Histopathology, Cellular Microscopy, Agriculture, Microbiology) with ~220K images, 619 categories, and 15M instances, claiming superior multi-domain performance and generalization over existing open-world counting methods.

Significance. If the cross-domain results hold without benchmark artifacts, the work would provide a useful generalist baseline for text-conditioned counting, unifying previously fragmented domain-specific tasks. The point-centric supervision strategy that accommodates heterogeneous annotations (point, box, density) and the release of code are concrete strengths that support reproducibility and further research.

major comments (2)
  1. [CLOC construction / dataset section] CLOC construction (described in the abstract and the dataset section): the manuscript provides no analysis demonstrating absence of overlapping images/near-duplicates across the six domains, category-name collisions with inconsistent definitions, or annotation-style leakage (point vs. box vs. density maps) from the source collections. Because the headline claim of multi-domain generalization rests on CLOC being an unbiased test set, this omission is load-bearing.
  2. [Experiments section] Experiments section: the reported outperformance is presented without error bars, ablation tables isolating the contribution of each counter or the fusion step, or statistics on domain balance and category distribution within CLOC. This makes it impossible to assess whether the gains are robust or sensitive to post-hoc choices.
minor comments (2)
  1. [Abstract / Method] The abstract states the fusion is 'parameter-free' but does not explicitly define the fusion rule or show that no learned weights are involved; a short equation or pseudocode would clarify this.
  2. [Method] Notation for the two counters (Region-level Sparse Counter, Pixel-level Dense Counter) is introduced without a table comparing their supervision signals or output formats; a small comparison table would improve readability.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive feedback and positive assessment of the work's potential significance. We address the two major comments below and commit to revisions that strengthen the manuscript without altering its core claims.

read point-by-point responses
  1. Referee: [CLOC construction / dataset section] CLOC construction (described in the abstract and the dataset section): the manuscript provides no analysis demonstrating absence of overlapping images/near-duplicates across the six domains, category-name collisions with inconsistent definitions, or annotation-style leakage (point vs. box vs. density maps) from the source collections. Because the headline claim of multi-domain generalization rests on CLOC being an unbiased test set, this omission is load-bearing.

    Authors: We agree that explicit verification of cross-domain image overlaps, near-duplicates, category-name consistency, and annotation-style leakage is necessary to support the multi-domain generalization claims. The current manuscript does not include such analysis. In the revision we will add a dedicated subsection reporting: (i) image-level deduplication checks (e.g., perceptual hashing and embedding similarity thresholds) across the six source collections, (ii) manual review of category-name collisions with harmonized definitions where needed, and (iii) confirmation that annotation formats were converted without leakage of supervision style into the evaluation splits. If any overlaps are found, we will document removal statistics. revision: yes

  2. Referee: [Experiments section] Experiments section: the reported outperformance is presented without error bars, ablation tables isolating the contribution of each counter or the fusion step, or statistics on domain balance and category distribution within CLOC. This makes it impossible to assess whether the gains are robust or sensitive to post-hoc choices.

    Authors: We concur that the absence of error bars, component-wise ablations, and dataset statistics limits interpretability of the reported gains. The revision will include: (i) standard deviations or confidence intervals over multiple random seeds for all main results, (ii) ablation tables that isolate the Region-level Sparse Counter, Pixel-level Dense Counter, and Complementary Count Fusion, and (iii) supplementary tables/figures showing per-domain image counts, category distributions, and instance-density histograms within CLOC. These additions will be placed in the Experiments section and appendix. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; claims rest on empirical evaluation of reorganized benchmark

full rationale

The paper's central claims rest on constructing CLOC by reorganizing existing public datasets and then reporting experimental accuracy of the proposed dual-granularity counters on that benchmark. No derivation step reduces a prediction or first-principles result to its own inputs by construction, no fitted parameter is relabeled as a prediction, and no load-bearing uniqueness theorem or ansatz is imported via self-citation. The performance numbers are obtained from standard train/test splits on the assembled data rather than being forced by the model definition itself.

Assumptions & free parameters 0 free parameters · 0 assumptions · 2 invented entities

Abstract-only review yields no explicit free parameters, axioms or invented entities beyond the two named counters introduced as model components; no independent evidence is supplied for those components.

invented entities (2)
  • Region-level Sparse Counter
    purpose: Provides object-level anchors for large and sparse targets
    Component introduced in the model description.
  • Pixel-level Dense Counter
    purpose: Handles small, crowded, and weakly bounded targets via dense point prediction
    Component introduced in the model description.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Count Anything." pith.science (2026). https://pith.science/paper/U4XB4HSG

@misc{pith2026260530846,
  author       = {Pith},
  title        = {Pith review of: Count Anything},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/U4XB4HSG}},
  note         = {Machine review of arXiv:2605.30846}
}
read the original abstract

Object counting remains fragmented across domain-specific datasets and task formulations, despite rapid progress in generalist vision models. Existing counting models are often tailored to scenarios such as crowds, vehicles, cells, crops, or remote-sensing objects, and thus struggle to generalize across categories, visual domains, object scales, and density distributions. In this paper, we study text-guided object counting across domains, where a model takes an image and a natural-language query as input and returns an instance-grounded set of target points whose cardinality gives the count. This formulation unifies category-conditioned counting with interpretable spatial localization. To support this setting, we construct CLOC, a Cross-domain Large-scale Object Counting dataset that reorganizes diverse public data sources into a unified benchmark. CLOC covers six visual domains: General Scene, Remote Sensing, Histopathology, Cellular Microscopy, Agriculture, and Microbiology, with about 220K images, 619 categories, and 15M object instances. Based on CLOC, we propose Count Anything, a generalist model for text-guided object counting. Unlike density-map-based methods, which dominate counting models, Count Anything adopts discrete instance points and performs dual-granularity instance enumeration. A Region-level Sparse Counter provides object-level anchors for large and sparse targets, while a Pixel-level Dense Counter handles small, crowded, and weakly bounded targets via dense point prediction. A point-centric supervision strategy enables learning from heterogeneous annotations, and Complementary Count Fusion combines both counters in a parameter-free manner. Extensive experiments show that Count Anything achieves strong accuracy and multi-domain generalization, outperforming existing open-world counting methods. Code is available at: https://github.com/Mengqi-Lei/count-anything.

Figures

Figures reproduced from arXiv: 2605.30846 by the authors.

Figure 1
Figure 1. Overall framework of the proposed Count Anything. [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. Illustration of CCF. Each RSC prediction is compared with the nearest PDC point inside its region, and only the higher-confidence one is kept. During inference, RSC and PDC produce region-level candidates and dense point candidates, respectively. Di￾rect concatenation may double-count clearly bounded tar￾gets predicted by both branches, while removing all PDC points inside each RSC region may discard multiple true i… view at source ↗
Figure 3
Figure 3. Qualitative comparison of counting predictions. Text prompts are shown on the left, and numbers denote ground-truth or predicted counts [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (14 more)
Figure 4
Figure 4. Figure 4: Effect of training data scale. Data scaling [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Qualitative visualization of Complementary Count Fusion. Orange boxes and points denote [PITH_FULL_IMAGE:figures/full_fig_p021_5.png]
Figure 6
Figure 6. Figure 6: Additional qualitative comparisons on the General Scene domain. [PITH_FULL_IMAGE:figures/full_fig_p022_6.png]
Figure 7
Figure 7. Figure 7: Additional qualitative comparisons on the Remote Sensing domain. [PITH_FULL_IMAGE:figures/full_fig_p023_7.png]
Figure 8
Figure 8. Figure 8: Additional qualitative comparisons on the Histopathology domain. [PITH_FULL_IMAGE:figures/full_fig_p024_8.png]
Figure 9
Figure 9. Figure 9: Additional qualitative comparisons on the Cellular Microscopy domain. [PITH_FULL_IMAGE:figures/full_fig_p025_9.png]
Figure 10
Figure 10. Figure 10: Additional qualitative comparisons on the Agriculture domain. [PITH_FULL_IMAGE:figures/full_fig_p026_10.png]
Figure 11
Figure 11. Figure 11: Additional qualitative comparisons on the Microbiology domain. [PITH_FULL_IMAGE:figures/full_fig_p027_11.png]
Figure 12
Figure 12. Figure 12: Representative visual examples from the six visual domains of CLOC. Each row shows [PITH_FULL_IMAGE:figures/full_fig_p030_12.png]
Figure 13
Figure 13. Figure 13: Category-level image frequency distribution. The figure shows the top 80 categories [PITH_FULL_IMAGE:figures/full_fig_p031_13.png]
Figure 14
Figure 14. Figure 14: Target-count distribution. The x-axis indicates the target-count range for each image [PITH_FULL_IMAGE:figures/full_fig_p032_14.png]
Figure 15
Figure 15. Figure 15: Overview of the CLOC construction pipeline. CLOC starts from multi-source raw datasets [PITH_FULL_IMAGE:figures/full_fig_p033_15.png]
Figure 16
Figure 16. Figure 16: Examples of source-specific special annotations requiring countability audit before [PITH_FULL_IMAGE:figures/full_fig_p034_16.png]
Figure 17
Figure 17. Figure 17: Distribution of original and derived samples across target-count ranges. The figure shows [PITH_FULL_IMAGE:figures/full_fig_p039_17.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

105 extracted references · 6 canonical work pages

  1. [1]

    Foundation models defining a new era in vision: A survey and outlook.IEEE Trans

    Muhammad Awais, Muzammal Naseer, Salman Khan, Rao Muhammad Anwer, Hisham Cholakkal, Mubarak Shah, Ming-Hsuan Yang, and Fahad Shahbaz Khan. Foundation models defining a new era in vision: A survey and outlook.IEEE Trans. Pattern Anal. Mach. Intell., 47(4):2245–2264, 2025

  2. [2]

    Vision foundation models in remote sensing: A survey.IEEE Geosci

    Siqi Lu, Junlin Guo, James R Zimmer-Dauphinee, Jordan M Nieusma, Xiao Wang, Parker VanValkenburgh, Steven A Wernke, and Yuankai Huo. Vision foundation models in remote sensing: A survey.IEEE Geosci. Remote Sens. Mag., 13(3):190–215, 2025

  3. [3]

    Florence: A New Foundation Model for Computer Vision

    Lu Yuan, Dongdong Chen, Yi-Ling Chen, Noel Codella, Xiyang Dai, Jianfeng Gao, Houdong Hu, Xuedong Huang, Boxin Li, Chunyuan Li, et al. Florence: A new foundation model for computer vision.arXiv preprint arXiv:2111.11432, 2021

  4. [4]

    VLCounter: Text-aware visual representa- tion for zero-shot object counting

    Seunggu Kang, WonJun Moon, Euiyeon Kim, and Jae-Pil Heo. VLCounter: Text-aware visual representa- tion for zero-shot object counting. InAAAI, volume 38, pages 2714–2722, 2024

  5. [5]

    Zero-Shot Object Counting With Good Exemplars

    Huilin Zhu, Jingling Yuan, Zhengwei Yang, Yu Guo, Zheng Wang, Xian Zhong, and Shengfeng He. Zero-Shot Object Counting With Good Exemplars. InEur . Conf. Comput. Vis., pages 368–385, 2024

  6. [6]

    CountGD: Multi-modal open-world counting

    Niki Amini-Naieni, Tengda Han, and Andrew Zisserman. CountGD: Multi-modal open-world counting. In Adv. Neural Inform. Process. Syst., volume 37, pages 48810–48837, 2024

  7. [7]

    CountGD++: Generalized prompting for open-world counting

    Niki Amini-Naieni and Andrew Zisserman. CountGD++: Generalized prompting for open-world counting. InIEEE Conf. Comput. Vis. Pattern Recog., 2026. 10

  8. [8]

    Deep learning in crowd counting: A survey.CAAI Trans

    Lijia Deng, Qinghua Zhou, Shuihua Wang, Juan Manuel Górriz, and Yudong Zhang. Deep learning in crowd counting: A survey.CAAI Trans. on Intell. Tech., 9(5):1043–1077, 2024

Show all 105 references
  1. [9]

    Zero-shot object counting

    Jingyi Xu, Hieu Le, Vu Nguyen, Viresh Ranjan, and Dimitris Samaras. Zero-shot object counting. InIEEE Conf. Comput. Vis. Pattern Recog., pages 15548–15557, 2023

  2. [10]

    Rethinking counting and localization in crowds: A purely point-based framework

    Qingyu Song, Changan Wang, Zhengkai Jiang, Yabiao Wang, Ying Tai, Chengjie Wang, Jilin Li, Feiyue Huang, and Yang Wu. Rethinking counting and localization in crowds: A purely point-based framework. InInt. Conf. Comput. Vis., pages 3365–3374, 2021

  3. [11]

    Learning To Count Anything: Reference-less class-agnostic counting with weak supervision

    Michael Hobley and Victor Prisacariu. Learning To Count Anything: Reference-less class-agnostic counting with weak supervision. InIEEE Conf. Comput. Vis. Pattern Recog. Worksh., 2023

  4. [12]

    Learning To Count Everything

    Viresh Ranjan, Udbhav Sharma, Thu Nguyen, and Minh Hoai. Learning To Count Everything. InIEEE Conf. Comput. Vis. Pattern Recog., pages 3394–3403, 2021

  5. [13]

    Few-shot Object Counting and Detection

    Thanh Nguyen, Chau Pham, Khoi Nguyen, and Minh Hoai. Few-shot Object Counting and Detection. In Eur . Conf. Comput. Vis., pages 348–365, 2022

  6. [14]

    Single-image crowd counting via multi-column convolutional neural network

    Yingying Zhang, Desen Zhou, Siqin Chen, Shenghua Gao, and Yi Ma. Single-image crowd counting via multi-column convolutional neural network. InIEEE Conf. Comput. Vis. Pattern Recog., pages 589–597, 2016

  7. [15]

    Improving Point-Based Crowd Counting and Localization Based on Auxiliary Point Guidance

    I-Hsiang Chen, Wei-Ting Chen, Yu-Wei Liu, Ming-Hsuan Yang, and Sy-Yen Kuo. Improving Point-Based Crowd Counting and Localization Based on Auxiliary Point Guidance. InEur . Conf. Comput. Vis., pages 428–444, 2024

  8. [16]

    CCTrans: Simplifying and improving crowd counting with transformer.arXiv preprint arXiv:2109.14483, 2021

    Ye Tian, Xiangxiang Chu, and Hongpeng Wang. CCTrans: Simplifying and improving crowd counting with transformer.arXiv preprint arXiv:2109.14483, 2021

  9. [17]

    Microsoft COCO: Common objects in context

    Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft COCO: Common objects in context. InEur . Conf. Comput. Vis., pages 740–755, 2014

  10. [18]

    Mark Everingham, Luc Van Gool, Christopher K. I. Williams, John Winn, and Andrew Zisserman. The PASCAL Visual Object Classes (VOC) Challenge.Int. J. Comput. Vis., 88(2):303–338, 2010

  11. [19]

    Sindagi, Rajeev Yasarla, and Vishal M

    Vishwanath A. Sindagi, Rajeev Yasarla, and Vishal M. Patel. JHU-CROWD++: Large-scale crowd counting dataset and a benchmark method.IEEE Trans. Pattern Anal. Mach. Intell., 44(5):2594–2609, 2022

  12. [20]

    NuInsSeg: A fully annotated dataset for nuclei instance segmentation in H&E-stained histological images.Sci

    Amirreza Mahbod, Christine Polak, Katharina Feldmann, Rumsha Khan, Katharina Gelles, Georg Dorffner, Ramona Woitek, Sepideh Hatamikia, and Isabella Ellinger. NuInsSeg: A fully annotated dataset for nuclei instance segmentation in H&E-stained histological images.Sci. Data, 11(1...

  13. [21]

    A Dataset and a Technique for Generalized Nuclear Segmentation for Computational Pathology.IEEE Trans

    Neeraj Kumar, Ruchika Verma, Sanuj Sharma, Surabhi Bhargava, Abhishek Vahadane, and Amit Sethi. A Dataset and a Technique for Generalized Nuclear Segmentation for Computational Pathology.IEEE Trans. Med. Imag., 36(7):1550–1560, 2017

  14. [22]

    CellBinDB: A large-scale multimodal annotated dataset for cell segmentation with benchmarking of universal models.GigaScience, 14:giaf069, 2025

    Can Shi, Jinghong Fan, Zhonghan Deng, Huanlin Liu, Qiang Kang, Yumei Li, Jing Guo, Jingwen Wang, Jinjiang Gong, Sha Liao, Ao Chen, Ying Zhang, and Mei Li. CellBinDB: A large-scale multimodal annotated dataset for cell segmentation with benchmarking of universal models.GigaScie...

  15. [23]

    Objects365: A large-scale, high-quality dataset for object detection

    Shuai Shao, Zeming Li, Tianyuan Zhang, Chao Peng, Gang Yu, Xiangyu Zhang, Jing Li, and Jian Sun. Objects365: A large-scale, high-quality dataset for object detection. InInt. Conf. Comput. Vis., pages 8430–8439, 2019

  16. [24]

    Detection and Tracking Meet Drones Challenge.IEEE Trans

    Pengfei Zhu, Longyin Wen, Dawei Du, Xiao Bian, Heng Fan, Qinghua Hu, and Haibin Ling. Detection and Tracking Meet Drones Challenge.IEEE Trans. Pattern Anal. Mach. Intell., 44(11):7380–7399, 2022

  17. [25]

    Ahmed Raza, Nasir M

    Ruchika Verma, Neeraj Kumar, Abhijeet Patil, Nikhil Cherian Kurian, Swapnil Rane, Simon Graham, Quoc Dang Vu, Mieke Zwager, Shan E. Ahmed Raza, Nasir M. Rajpoot, et al. MoNuSAC2020: A multi-organ nuclei segmentation and classification challenge.IEEE Trans. Med. Imag., 40(12):3...

  18. [26]

    Venkatesh Babu

    Deepak Babu Sam, Skand Vishwanath Peri, Mukuntha Narayanan Sundararaman, Amogh Kamath, and R. Venkatesh Babu. Locate, size, and count: Accurately resolving people in dense crowds via detection. IEEE Trans. Pattern Anal. Mach. Intell., 43(8):2739–2751, 2021. 11

  19. [27]

    Meng-Ru Hsieh, Yen-Liang Lin, and Winston H. Hsu. Drone-Based Object Counting by Spatially Regularized Regional Proposal Networks. InInt. Conf. Comput. Vis., pages 4145–4153, 2017

  20. [28]

    Distribution Matching for Crowd Counting

    Boyu Wang, Huidong Liu, Dimitris Samaras, and Minh Hoai Nguyen. Distribution Matching for Crowd Counting. InAdv. Neural Inform. Process. Syst., pages 1595–1607, 2020

  21. [29]

    CLIP-Count: Towards text-guided zero-shot object counting

    Ruixiang Jiang, Lingbo Liu, and Changwen Chen. CLIP-Count: Towards text-guided zero-shot object counting. InACM Int. Conf. Multimedia, pages 4535–4545, 2023

  22. [30]

    Consistency-aware anchor pyramid network for crowd localization.IEEE Trans

    Xinyan Liu, Guorong Li, Yuankai Qi, Zhenjun Han, Anton van den Hengel, Nicu Sebe, Ming-Hsuan Yang, and Qingming Huang. Consistency-aware anchor pyramid network for crowd localization.IEEE Trans. Pattern Anal. Mach. Intell., 2024

  23. [31]

    Kevin Zhou, and Jie Chen

    Zhongyi Huang, Yao Ding, Guoli Song, Lin Wang, Ruizhe Geng, Hongliang He, Shan Du, Xia Liu, Yonghong Tian, Yongsheng Liang, S. Kevin Zhou, and Jie Chen. BCData: A large-scale dataset and benchmark for cell detection and counting. InMed. Image Comput. Comput. Assist. Interv., p...

  24. [32]

    TasselNet: Counting maize tassels in the wild via local counts regression network.Plant Methods, 13(1):79, 2017

    Hao Lu, Zhiguo Cao, Yang Xiao, Bohan Zhuang, and Chunhua Shen. TasselNet: Counting maize tassels in the wild via local counts regression network.Plant Methods, 13(1):79, 2017

  25. [33]

    LVIS: A dataset for large vocabulary instance segmentation

    Agrim Gupta, Piotr Dollár, and Ross Girshick. LVIS: A dataset for large vocabulary instance segmentation. InIEEE Conf. Comput. Vis. Pattern Recog., pages 5356–5364, 2019

  26. [34]

    Plummer, Liwei Wang, Chris M

    Bryan A. Plummer, Liwei Wang, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, and Svetlana Lazebnik. Flickr30k Entities: Collecting region-to-phrase correspondences for richer image-to-sentence models. InInt. Conf. Comput. Vis., pages 2641–2649, 2015

  27. [35]

    SAM 3: Segment anything with concepts

    Nicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath, Ronghang Hu, Didac Suris, Chaitanya Ryali, Kalyan Vasudev Alwala, Haitham Khedr, Andrew Huang, et al. SAM 3: Segment anything with concepts. InInt. Conf. Learn. Represent., 2026

  28. [36]

    End-to-End Object Detection with Transformers

    Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-End Object Detection with Transformers. InEur . Conf. Comput. Vis., pages 213–229, 2020

  29. [37]

    A Low-Shot Object Counting Network With Iterative Prototype Adaptation

    Nikola Ðuki´c, Alan Lukežiˇc, Vitjan Zavrtanik, and Matej Kristan. A Low-Shot Object Counting Network With Iterative Prototype Adaptation. InInt. Conf. Comput. Vis., pages 18872–18881, 2023

  30. [38]

    Open-World Text-Specified Object Counting

    Niki Amini-Naieni, Kiana Amini-Naieni, Tengda Han, and Andrew Zisserman. Open-World Text-Specified Object Counting. InBrit. Mach. Vis. Conf., 2023

  31. [39]

    YOLO-Count: Differentiable object counting for text-to-image generation

    Guanning Zeng, Xiang Zhang, Zirui Wang, Haiyang Xu, Zeyuan Chen, Bingnan Li, and Zhuowen Tu. YOLO-Count: Differentiable object counting for text-to-image generation. InInt. Conf. Comput. Vis., pages 16765–16775, 2025

  32. [40]

    Yifei Qian, Zhongliang Guo, Bowen Deng, Chun Tong Lei, Shuai Zhao, Chun Pong Lau, Xiaopeng Hong, and Michael P. Pound. T2ICount: Enhancing cross-modal understanding for zero-shot counting. InIEEE Conf. Comput. Vis. Pattern Recog., pages 25336–25345, 2025

  33. [41]

    CountSE: Soft exemplar open-set object counting

    Shuai Liu, Peng Zhang, Shiwei Zhang, and Wei Ke. CountSE: Soft exemplar open-set object counting. In Int. Conf. Comput. Vis., pages 21536–21546, 2025

  34. [42]

    Grounded Language-Image Pre-Training

    Liunian Harold Li, Pengchuan Zhang, Haotian Zhang, Jianwei Yang, Chunyuan Li, Yiwu Zhong, Lijuan Wang, Lu Yuan, Lei Zhang, Jenq-Neng Hwang, et al. Grounded Language-Image Pre-Training. InIEEE Conf. Comput. Vis. Pattern Recog., pages 10965–10975, 2022

  35. [43]

    YOLO-World: Real- time open-vocabulary object detection

    Tianheng Cheng, Lin Song, Yixiao Ge, Wenyu Liu, Xinggang Wang, and Ying Shan. YOLO-World: Real- time open-vocabulary object detection. InIEEE Conf. Comput. Vis. Pattern Recog., pages 16901–16911, 2024

  36. [44]

    Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Qing Jiang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, and Lei Zhang. Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection. InEur . Conf. Comput. Vis., pages 38–55, 2024

  37. [45]

    YOLOE: Real-time seeing anything

    Ao Wang, Lihao Liu, Hui Chen, Zijia Lin, Jungong Han, and Guiguang Ding. YOLOE: Real-time seeing anything. InInt. Conf. Comput. Vis., pages 24591–24602, 2025. 12

  38. [46]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. InInt. Conf. Learn. Represent., 2022

  39. [47]

    Decoupled weight decay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. InInt. Conf. Learn. Represent., 2019

  40. [48]

    Redesigning Multi-Scale Neural Network for Crowd Counting.IEEE Trans

    Zhipeng Du, Miaojing Shi, Jiankang Deng, and Stefanos Zafeiriou. Redesigning Multi-Scale Neural Network for Crowd Counting.IEEE Trans. Image Process., 32:3664–3678, 2023

  41. [49]

    CSRNet: Dilated convolutional neural networks for understanding the highly congested scenes

    Yuhong Li, Xiaofan Zhang, and Deming Chen. CSRNet: Dilated convolutional neural networks for understanding the highly congested scenes. InIEEE Conf. Comput. Vis. Pattern Recog., pages 1091–1100, 2018

  42. [50]

    Scale Aggregation Network for Accurate and Efficient Crowd Counting

    Xinkun Cao, Zhipeng Wang, Yanyun Zhao, and Fei Su. Scale Aggregation Network for Accurate and Efficient Crowd Counting. InEur . Conf. Comput. Vis., pages 734–750, 2018

  43. [51]

    Bayesian loss for crowd count estimation with point supervision

    Zhiheng Ma, Xing Wei, Xiaopeng Hong, and Yihong Gong. Bayesian loss for crowd count estimation with point supervision. InInt. Conf. Comput. Vis., pages 6142–6151, 2019

  44. [52]

    CounTR: Transformer-based generalised visual counting

    Chang Liu, Yujie Zhong, Andrew Zisserman, and Weidi Xie. CounTR: Transformer-based generalised visual counting. InBrit. Mach. Vis. Conf., 2022

  45. [53]

    Point, Segment and Count: A generalized framework for object counting

    Zhizhong Huang, Mingliang Dai, Yi Zhang, Junping Zhang, and Hongming Shan. Point, Segment and Count: A generalized framework for object counting. InIEEE Conf. Comput. Vis. Pattern Recog., pages 17067–17076, 2024

  46. [54]

    DA VE: A detect-and-verify paradigm for low-shot counting

    Jer Pelhan, Alan Lukežiˇc, Vitjan Zavrtanik, and Matej Kristan. DA VE: A detect-and-verify paradigm for low-shot counting. InIEEE Conf. Comput. Vis. Pattern Recog., pages 23293–23302, 2024

  47. [55]

    Open-World Object Counting in Videos

    Niki Amini-Naieni and Andrew Zisserman. Open-World Object Counting in Videos. InAAAI, volume 40, pages 2300–2308, 2026

  48. [56]

    NWPU-Crowd: A large-scale benchmark for crowd counting and localization.IEEE Trans

    Qi Wang, Junyu Gao, Wei Lin, and Xuelong Li. NWPU-Crowd: A large-scale benchmark for crowd counting and localization.IEEE Trans. Pattern Anal. Mach. Intell., 43(6):2141–2149, 2021

  49. [57]

    Badhon, Curtis Pozniak, Benoit de Solan, Andreas Hund, Scott C

    Etienne David, Simon Madec, Pouria Sadeghi-Tehran, Helge Aasen, Bangyou Zheng, Shouyang Liu, Norbert Kirchgessner, Goro Ishikawa, Koichi Nagasawa, Minhajul A. Badhon, Curtis Pozniak, Benoit de Solan, Andreas Hund, Scott C. Chapman, Frédéric Baret, Ian Stavness, and Wei Guo. Gl...

  50. [58]

    AGAR: A microbial colony dataset for deep learning detection, 2021

    Sylwia Majchrowska, Jarosław Pawłowski, Grzegorz Guła, Tomasz Bonus, Agata Hanas, Adam Loch, Agnieszka Pawlak, Justyna Roszkowiak, Tomasz Golan, and Zuzanna Drulis-Kawa. AGAR: A microbial colony dataset for deep learning detection, 2021

  51. [59]

    OmniCount: Multi-label object counting with semantic-geometric priors

    Anindya Mondal, Sauradip Nag, Xiatian Zhu, and Anjan Dutta. OmniCount: Multi-label object counting with semantic-geometric priors. InAAAI, volume 39, pages 19537–19545, 2025

  52. [60]

    DOTA: A large-scale dataset for object detection in aerial images

    Gui-Song Xia, Xiang Bai, Jian Ding, Zhen Zhu, Serge Belongie, Jiebo Luo, Mihai Datcu, Marcello Pelillo, and Liangpei Zhang. DOTA: A large-scale dataset for object detection in aerial images. InIEEE Conf. Comput. Vis. Pattern Recog., pages 3974–3983, 2018

  53. [61]

    xView: Objects in context in overhead imagery.arXiv preprint arXiv:1802.07856, 2018

    Darius Lam, Richard Kuzma, Kevin McGee, Samuel Dooley, Michael Laielli, Matthew Klaric, Yaroslav Bulatov, and Brendan McCord. xView: Objects in context in overhead imagery.arXiv preprint arXiv:1802.07856, 2018

  54. [62]

    The Cityscapes Dataset for Semantic Urban Scene Understanding

    Marius Cordts, Mohamed Omran, Sebastian Ramos, Timo Rehfeld, Markus Enzweiler, Rodrigo Benen- son, Uwe Franke, Stefan Roth, and Bernt Schiele. The Cityscapes Dataset for Semantic Urban Scene Understanding. InIEEE Conf. Comput. Vis. Pattern Recog., pages 3213–3223, 2016

  55. [63]

    Jackson, Nabeel Khalid, Nicola Bevan, Timothy Dale, Andreas Dengel, Sheraz Ahmed, Johan Trygg, and Rickard Sjögren

    Christoffer Edlund, Timothy R. Jackson, Nabeel Khalid, Nicola Bevan, Timothy Dale, Andreas Dengel, Sheraz Ahmed, Johan Trygg, and Rickard Sjögren. LIVECell: A large-scale dataset for label-free live cell segmentation.Nat. Methods, 18(9):1038–1045, 2021

  56. [64]

    SoybeanNet: Transformer-based convolutional neural network for soybean pod counting from unmanned aerial vehicle (UA V) images.Comput

    Jiajia Li, Raju Thada Magar, Dong Chen, Feng Lin, Dechun Wang, Xiang Yin, Weichao Zhuang, and Zhaojian Li. SoybeanNet: Transformer-based convolutional neural network for soybean pod counting from unmanned aerial vehicle (UA V) images.Comput. Electron. Agric., 220:108861, 2024. 13

  57. [65]

    Learning Transferable Visual Models From Natural Language Supervision

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning Transferable Visual Models From Natural Language Supervision. InInt. Conf. Mach. L...

  58. [66]

    Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick. Segment Anything. InInt. Conf. Comput. Vis., pages 4015–4026, 2023

  59. [67]

    MedMNIST v2: A large-scale lightweight benchmark for 2D and 3D biomedical image classification

    Jiancheng Yang, Rui Shi, Donglai Wei, Zequan Liu, Lin Zhao, Bilian Ke, Hanspeter Pfister, and Bingbing Ni. MedMNIST v2: A large-scale lightweight benchmark for 2D and 3D biomedical image classification. Sci. Data, 10(1):41, 2023

  60. [68]

    Landman, Geert Litjens, Bjoern Menze, Olaf Ronneberger, Ronald M

    Michela Antonelli, Annika Reinke, Spyridon Bakas, Keyvan Farahani, Annette Kopp-Schneider, Bennett A. Landman, Geert Litjens, Bjoern Menze, Olaf Ronneberger, Ronald M. Summers, Bram van Ginneken, et al. The Medical Segmentation Decathlon.Nat. Commun., 13(1):4128, 2022

  61. [69]

    MSeg: A composite dataset for multi-domain semantic segmentation

    John Lambert, Zhuang Liu, Ozan Sener, James Hays, and Vladlen Koltun. MSeg: A composite dataset for multi-domain semantic segmentation. InIEEE Conf. Comput. Vis. Pattern Recog., pages 2879–2888, 2020

  62. [70]

    BigDetection: A large-scale benchmark for improved object detector pre-training

    Likun Cai, Zhi Zhang, Yi Zhu, Li Zhang, Mu Li, and Xiangyang Xue. BigDetection: A large-scale benchmark for improved object detector pre-training. InIEEE Conf. Comput. Vis. Pattern Recog. Worksh., pages 4777–4787, 2022

  63. [71]

    rebar counting dataset

    fyp2. rebar counting dataset. https://universe.roboflow.com/fyp2-czq30/ rebar-counting-12vha, September 2024. visited on 2026-05-02

  64. [72]

    Paulo R. L. de Almeida, Luiz S. Oliveira, Alceu S. Britto Jr., Eunelson J. Silva Jr., and Alessandro L. Koerich. PKLot: A robust dataset for parking lot classification.Expert Syst. Appl., 42(11):4937–4949, 2015

  65. [73]

    Object Detection in Optical Remote Sensing Images: A survey and a new benchmark.ISPRS J

    Ke Li, Gang Wan, Gong Cheng, Liqiu Meng, and Junwei Han. Object Detection in Optical Remote Sensing Images: A survey and a new benchmark.ISPRS J. Photogramm. Remote Sens., 159:296–307, 2020

  66. [74]

    Object Detection in Aerial Images: A large-scale benchmark and challenges.IEEE Trans

    Jian Ding, Nan Xue, Gui-Song Xia, Xiang Bai, Wen Yang, Michael Ying Yang, Serge Belongie, Jiebo Luo, Mihai Datcu, Marcello Pelillo, and Liangpei Zhang. Object Detection in Aerial Images: A large-scale benchmark and challenges.IEEE Trans. Pattern Anal. Mach. Intell., 44(11):777...

  67. [75]

    Learning RoI Transformer for Oriented Object Detection in Aerial Images

    Jian Ding, Nan Xue, Yang Long, Gui-Song Xia, and Qikai Lu. Learning RoI Transformer for Oriented Object Detection in Aerial Images. InIEEE Conf. Comput. Vis. Pattern Recog., pages 2849–2858, 2019

  68. [76]

    Detection, Tracking, and Counting Meets Drones in Crowds: A benchmark

    Longyin Wen, Dawei Du, Pengfei Zhu, Qinghua Hu, Qilong Wang, Liefeng Bo, and Siwei Lyu. Detection, Tracking, and Counting Meets Drones in Crowds: A benchmark. InIEEE Conf. Comput. Vis. Pattern Recog., pages 7812–7821, 2021

  69. [77]

    NWPU-MOC: A benchmark for fine-grained multicategory object counting in aerial images.IEEE Trans

    Junyu Gao, Liangliang Zhao, and Xuelong Li. NWPU-MOC: A benchmark for fine-grained multicategory object counting in aerial images.IEEE Trans. Geosci. Remote Sens., 62:1–14, 2024

  70. [78]

    Multi-Class Geospatial Object Detection and Geographic Image Classification Based on Collection of Part Detectors.ISPRS J

    Gong Cheng, Junwei Han, Peicheng Zhou, and Lei Guo. Multi-Class Geospatial Object Detection and Geographic Image Classification Based on Collection of Part Detectors.ISPRS J. Photogramm. Remote Sens., 98:119–132, 2014

  71. [79]

    A Survey on Object Detection in Optical Remote Sensing Images.ISPRS J

    Gong Cheng and Junwei Han. A Survey on Object Detection in Optical Remote Sensing Images.ISPRS J. Photogramm. Remote Sens., 117:11–28, 2016

  72. [80]

    Learning Rotation-Invariant Convolutional Neural Networks for Object Detection in VHR Optical Remote Sensing Images.IEEE Trans

    Gong Cheng, Peicheng Zhou, and Junwei Han. Learning Rotation-Invariant Convolutional Neural Networks for Object Detection in VHR Optical Remote Sensing Images.IEEE Trans. Geosci. Remote Sens., 54(12):7405–7415, 2016

  73. [81]

    Accurate Object Localization in Remote Sensing Images Based on Convolutional Neural Networks.IEEE Trans

    Yang Long, Yiping Gong, Zhifeng Xiao, and Qing Liu. Accurate Object Localization in Remote Sensing Images Based on Convolutional Neural Networks.IEEE Trans. Geosci. Remote Sens., 55(5):2486–2498, 2017

  74. [82]

    Elliptic Fourier Transformation-Based Histograms of Oriented Gradients for Rotationally Invariant Object Detection in Remote-Sensing Images.Int

    Zhifeng Xiao, Qing Liu, Gefu Tang, and Xiaofang Zhai. Elliptic Fourier Transformation-Based Histograms of Oriented Gradients for Rotationally Invariant Object Detection in Remote-Sensing Images.Int. J. Remote Sens., 36(2):618–644, 2015

  75. [83]

    Enhancing People Localisation in Drone Imagery for Better Crowd Management by Utilising Every Pixel in High-Resolution Images.arXiv preprint arXiv:2502.04014, 2025

    Bartosz Ptak and Marek Kraft. Enhancing People Localisation in Drone Imagery for Better Crowd Management by Utilising Every Pixel in High-Resolution Images.arXiv preprint arXiv:2502.04014, 2025. 14

  76. [84]

    CoNIC Challenge: Pushing the frontiers of nuclear detection, segmentation, classification and counting.Med

    Simon Graham, Quoc Dang Vu, Mostafa Jahanifar, Martin Weigert, Uwe Schmidt, Wenhua Zhang, Jun Zhang, Sen Yang, Jinxi Xiang, Xiyue Wang, et al. CoNIC Challenge: Pushing the frontiers of nuclear detection, segmentation, classification and counting.Med. Image Anal., 92:103047, 2024

  77. [85]

    EndoNuke: Nuclei detection dataset for estrogen and progesterone stained IHC endometrium scans.Data, 7(6):75, 2022

    Anton Naumov, Egor Ushakov, Andrey Ivanov, Konstantin Midiber, Tatyana Khovanskaya, Alexandra Konyukova, Polina Vishnyakova, Sergei Nora, Liudmila Mikhaleva, Timur Fatkhudinov, and Evgeny Karpulevich. EndoNuke: Nuclei detection dataset for estrogen and progesterone stained IHC...

  78. [86]

    Lizard: A large-scale dataset for colonic nuclear instance segmentation and classification

    Simon Graham, Mostafa Jahanifar, Ayesha Azam, Mohammed Nimir, Yee-Wah Tsang, Katherine Dodd, Emily Hero, Harvir Sahota, Atisha Tank, Ksenija Benes, et al. Lizard: A large-scale dataset for colonic nuclear instance segmentation and classification. InInt. Conf. Comput. Vis. Work...

  79. [87]

    A Multi-Organ Nucleus Segmentation Challenge.IEEE Trans

    Neeraj Kumar, Ruchika Verma, Deepak Anand, Yanning Zhou, Omer Fahri Onder, Efstratios Tsougenis, Hao Chen, Pheng-Ann Heng, Jiahui Li, Zhiqiang Hu, et al. A Multi-Organ Nucleus Segmentation Challenge.IEEE Trans. Med. Imag., 39(5):1380–1391, 2020

  80. [88]

    Atteya, Hagar Hussein, Kareem Hosny Mohammed, Ehab Hafiz, Maha A

    Mohamed Amgad, Lamees A. Atteya, Hagar Hussein, Kareem Hosny Mohammed, Ehab Hafiz, Maha A. T. Elsebaie, Ahmed M. Alhusseiny, Mohamed Atef AlMoslemany, Abdelmagid M. Elmatboly, Philip A. Pappalardo, Rokia Adel Sakr, et al. NuCLS: A scalable crowdsourcing approach and dataset fo...

  81. [89]

    Lambert, and Bachir El Debs

    Mathieu Gendarme, Annika M. Lambert, and Bachir El Debs. BriFiSeg: A deep learning-based method for semantic and instance segmentation of nuclei in brightfield images.arXiv preprint arXiv:2211.03072, 2022

  82. [90]

    Learning To Count Objects in Images

    Victor Lempitsky and Andrew Zisserman. Learning To Count Objects in Images. InAdv. Neural Inform. Process. Syst., pages 1324–1332, 2010

  83. [91]

    Laine, Pedro M

    Christoph Spahn, Estibaliz Gómez-de Mariscal, Romain F. Laine, Pedro M. Pereira, Lucas von Chamier, Mia Conduit, Mariana G. Pinho, Guillaume Jacquemet, Séamus Holden, Mike Heilemann, et al. DeepBacs for multi-task bacterial image analysis using open-source deep learning approa...

  84. [92]

    Etienne David, Mario Serouart, Daniel Smith, Simon Madec, Kaaviya Velumani, Shouyang Liu, Xu Wang, Francisco Pinto, Shahameh Shafiee, Izzat S. A. Tahir, et al. Global Wheat Head Detection 2021: An improved dataset for benchmarking wheat head detection methods.Plant Phenomics, ...

  85. [93]

    people’s heads

    Luca Ciampi, Ali Azmoudeh, Elif Ecem Akbaba, Erdi Sarıta¸ s, Ziya Ata Yazıcı, Hazım Kemal Ekenel, Giuseppe Amato, and Fabrizio Falchi. A Survey on Class-Agnostic Counting: Advancements from reference-based to open-world text-guided approaches.Comput. Vis. Image Underst., 267:1...

  86. [94]

    Some source datasets adopt standard COCO-style JSON annotations, such as Objects365-2020, FSCD-LVIS, and LIVECell

    Parsing JSON-based annotation protocols.The original datasets vary substantially in their annotation protocols. Some source datasets adopt standard COCO-style JSON annotations, such as Objects365-2020, FSCD-LVIS, and LIVECell. These annotations are typically organized around i...

  87. [95]

    VOC-style datasets use XML files to record the target category, bounding box, difficult flag, and related fields for each image

    Parsing non-JSON annotation protocols.Beyond JSON annotations, we also process a variety of non-JSON annotation protocols. VOC-style datasets use XML files to record the target category, bounding box, difficult flag, and related fields for each image. Medical datasets such as ...

  88. [96]

    35439": {

    Unified instance representation.During conversion, different types of instance annotations are unified into two basic representations: point and bbox. For instances that already provide bounding boxes, we retain their axis-aligned bounding boxes and generate corresponding coun...

  89. [97]

    Public datasets often adopt different naming conventions, such as differences in capitalization, singular/plural forms, spaces, underscores, and special characters

    Category name normalization.We normalize category names so that categories from different data sources use consistent written forms. Public datasets often adopt different naming conventions, such as differences in capitalization, singular/plural forms, spaces, underscores, and...

  90. [98]

    one instance, one count

    Removing categories unsuitable for counting.We further remove categories that are not suitable as instance-level counting targets. Not all original categories in multi-source datasets can be stably associated with the “one instance, one count” assumption. Some categories have ...

  91. [99]

    For category consolidation in a multi-source dataset, it is not appropriate to simply promote all subclasses to their parent class

    Semantic category group construction.We construct semantic category groups for some fine- grained categories to support the later seen / unseen split. For category consolidation in a multi-source dataset, it is not appropriate to simply promote all subclasses to their parent c...

  92. [100]

    Since the dataset covers multiple visual domains, image resolutions vary substantially across data sources

    High-resolution image normalization.After category consolidation, some extremely high- resolution images need to be normalized in scale. Since the dataset covers multiple visual domains, image resolutions vary substantially across data sources. For example, DOTA [60, 74, 75], ...

  93. [101]

    Cropped sample generation.Cropped samples are generated by extracting local target regions from original images. Unlike simple random cropping, our cropping process prioritizes local windows whose target counts fall into predefined count ranges, thereby supplementing medium- a...

  94. [102]

    Stitched sample generation.Stitched sample generation also follows a target-count-oriented strategy. We first select candidate image-patch combinations according to predefined target-count ranges, so that the stitched samples can supplement specified medium- and high-count int...

  95. [103]

    In other words, cropped and stitched samples are retained only in the training set, while the validation and test sets do not contain any samples generated by cropping or stitching

    Derived-sample isolation.To avoid evaluation leakage introduced by data augmentation, we first construct the training, validation, and test sets at the original-sample level, and then generate cropped and stitched samples only from original images in the training set. In other...

  96. [104]

    To ensure stable cross-domain evaluation, the training, validation, and test sets all cover every visual domain

    Visual-domain proportion constraint.The split also explicitly considers the distribution of the six visual domains. To ensure stable cross-domain evaluation, the training, validation, and test sets all cover every visual domain. Under the constraints of derived-sample isolatio...

  97. [105]

    Existing category-conditioned counting datasets usually construct 39 unseen-category evaluation by ensuring that category names do not overlap

    Strict seen / unseen splitting based on category groups.For visual domains with rich category spaces, such as General Scene and Remote Sensing, we construct a seen / unseen evaluation setting to assess category generalization. Existing category-conditioned counting datasets us...

Pith tools

Reviewed June 28, 2026 · model on record in the stance chip above.