REVIEW 7 cited by
RP2K: A Large-Scale Retail Product Dataset for Fine-Grained Image Classification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce RP2K, a new large-scale retail product dataset for fine-grained image classification. Unlike previous datasets focusing on relatively few products, we collect more than 500,000 images of retail products on shelves belonging to 2000 different products. Our dataset aims to advance the research in retail object recognition, which has massive applications such as automatic shelf auditing and image-based product information retrieval. Our dataset enjoys following properties: (1) It is by far the largest scale dataset in terms of product categories. (2) All images are captured manually in physical retail stores with natural lightings, matching the scenario of real applications. (3) We provide rich annotations to each object, including the sizes, shapes and flavors/scents. We believe our dataset could benefit both computer vision research and retail industry. Our dataset is publicly available at https://www.pinlandata.com/rp2k_dataset.
Forward citations
Cited by 7 Pith papers
-
Instance-Level Generation for Representation Learning
Generating synthetic object instances from domain names and varying their backgrounds improves instance-level retrieval when used to fine-tune foundation models.
-
Illuminating Visual Identity in Universal Multimodal Embeddings
By adding identity-aware sampling and a contrastive loss on a new 28-dataset benchmark, the authors build multimodal embeddings that are far better at visual identity matching without losing general retrieval accuracy.
-
TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment
TIPSv2 improves dense patch-text alignment in vision-language pretraining through distillation and iBOT++ modifications, yielding models on par with or better than recent baselines on 9 tasks across 20 datasets.
-
RoboBenchMart: Benchmarking Robots in Retail Environment
RoboBenchMart is a new simulated retail benchmark showing that current generalist robot models perform poorly on common dark-store manipulation tasks.
-
Global-to-Local or Local-to-Global? Enhancing Image Retrieval with Efficient Local Search and Effective Global Re-ranking
A retrieval system that searches with local features and re-ranks with MDS-derived global embeddings sets new high scores on Revisited Oxford/Paris +1M, but is not state-of-the-art on base Oxford.
-
What Matters for Grocery Product Retrieval with Open Source Vision Language Models
Systematic zero-shot benchmarking of open-source VLMs on multimodal grocery product retrieval shows data quality outperforms scale, introduces semantic power density as an efficiency metric, and identifies a persisten...
-
A Mixed Diet Makes DINO An Omnivorous Vision Encoder
Fine-tuning the last blocks of DINOv2 with InfoNCE plus a teacher-anchoring loss maps RGB, depth, and segmentation views of the same scene to nearly identical feature vectors.
Discussion (0). Sign in to comment.