Pith. sign in

REVIEW 1 cited by

Object Detection in Indian Food Platters using Transfer Learning with YOLOv4

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.04841 v1 pith:GFWQF45P submitted 2022-05-10 cs.CV

Object Detection in Indian Food Platters using Transfer Learning with YOLOv4

classification cs.CV
keywords foodindiandishesobjectclassclassescontainsdataset-
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Object detection is a well-known problem in computer vision. Despite this, its usage and pervasiveness in the traditional Indian food dishes has been limited. Particularly, recognizing Indian food dishes present in a single photo is challenging due to three reasons: 1. Lack of annotated Indian food datasets 2. Non-distinct boundaries between the dishes 3. High intra-class variation. We solve these issues by providing a comprehensively labelled Indian food dataset- IndianFood10, which contains 10 food classes that appear frequently in a staple Indian meal and using transfer learning with YOLOv4 object detector model. Our model is able to achieve an overall mAP score of 91.8% and f1-score of 0.90 for our 10 class dataset. We also provide an extension of our 10 class dataset- IndianFood20, which contains 10 more traditional Indian food classes.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. DishSeg24k: A Large-Scale Benchmark for Food Segmentation with Stochastic Expert Decoding

    cs.CV 2026-07 conditional novelty 6.0

    DishSeg24k is a 24k-image dish-level food segmentation benchmark, and the FEAST model reports +3.21 mIoU over prior methods, mostly from its mixture-of-experts decoder.