Pith. sign in

REVIEW 7 cited by

RP2K: A Large-Scale Retail Product Dataset for Fine-Grained Image Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2006.12634 v7 pith:OXW7X5OE submitted 2020-06-22 cs.CV

classification cs.CV
keywords datasetretailproductproductsrp2kapplicationsclassificationfine-grained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce RP2K, a new large-scale retail product dataset for fine-grained image classification. Unlike previous datasets focusing on relatively few products, we collect more than 500,000 images of retail products on shelves belonging to 2000 different products. Our dataset aims to advance the research in retail object recognition, which has massive applications such as automatic shelf auditing and image-based product information retrieval. Our dataset enjoys following properties: (1) It is by far the largest scale dataset in terms of product categories. (2) All images are captured manually in physical retail stores with natural lightings, matching the scenario of real applications. (3) We provide rich annotations to each object, including the sizes, shapes and flavors/scents. We believe our dataset could benefit both computer vision research and retail industry. Our dataset is publicly available at https://www.pinlandata.com/rp2k_dataset.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Instance-Level Generation for Representation Learning

    cs.CV 2025-10 conditional novelty 7.0 of 10

    Generating synthetic object instances from domain names and varying their backgrounds improves instance-level retrieval when used to fine-tune foundation models.

  2. Illuminating Visual Identity in Universal Multimodal Embeddings

    cs.CV 2026-08 conditional novelty 6.0 of 10

    By adding identity-aware sampling and a contrastive loss on a new 28-dataset benchmark, the authors build multimodal embeddings that are far better at visual identity matching without losing general retrieval accuracy.

  3. TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    TIPSv2 improves dense patch-text alignment in vision-language pretraining through distillation and iBOT++ modifications, yielding models on par with or better than recent baselines on 9 tasks across 20 datasets.

  4. RoboBenchMart: Benchmarking Robots in Retail Environment

    cs.RO 2025-11 conditional novelty 6.0 of 10

    RoboBenchMart is a new simulated retail benchmark showing that current generalist robot models perform poorly on common dark-store manipulation tasks.

  5. Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking

    cs.IR 2025-09 conditional novelty 6.0 of 10

    A retrieval system that searches with local features and re-ranks with MDS-derived global embeddings sets new high scores on Revisited Oxford/Paris +1M, but is not state-of-the-art on base Oxford.

  6. What Matters for Grocery Product Retrieval with Open Source Vision Language Models

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    Systematic zero-shot benchmarking of open-source VLMs on multimodal grocery product retrieval shows data quality outperforms scale, introduces semantic power density as an efficiency metric, and identifies a persisten...

  7. A Mixed Diet Makes DINO An Omnivorous Vision Encoder

    cs.CV 2026-02 conditional novelty 5.0 of 10

    Fine-tuning the last blocks of DINOv2 with InfoNCE plus a teacher-anchoring loss maps RGB, depth, and segmentation views of the same scene to nearly identical feature vectors.

Pith tools