REVIEW 2 cited by
FoodSAM: Any Food Segmentation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we explore the zero-shot capability of the Segment Anything Model (SAM) for food image segmentation. To address the lack of class-specific information in SAM-generated masks, we propose a novel framework, called FoodSAM. This innovative approach integrates the coarse semantic mask with SAM-generated masks to enhance semantic segmentation quality. Besides, we recognize that the ingredients in food can be supposed as independent individuals, which motivated us to perform instance segmentation on food images. Furthermore, FoodSAM extends its zero-shot capability to encompass panoptic segmentation by incorporating an object detector, which renders FoodSAM to effectively capture non-food object information. Drawing inspiration from the recent success of promptable segmentation, we also extend FoodSAM to promptable segmentation, supporting various prompt variants. Consequently, FoodSAM emerges as an all-encompassing solution capable of segmenting food items at multiple levels of granularity. Remarkably, this pioneering framework stands as the first-ever work to achieve instance, panoptic, and promptable segmentation on food images. Extensive experiments demonstrate the feasibility and impressing performance of FoodSAM, validating SAM's potential as a prominent and influential tool within the domain of food image segmentation. We release our code at https://github.com/jamesjg/FoodSAM.
Forward citations
Cited by 2 Pith papers
-
Food Image Segmentation with LLM-Derived Ingredient Labels and Multimodal Fusion
Injecting BERT-encoded, LLM-generated ingredient labels into segmentation features and decoder queries raises FoodSeg103 mIoU from 51.9 (Mask2Former baseline) to 54.4 with LIM-F and 55.0 with LIM-Q.
-
Swin-TUNA : A Novel PEFT Approach for Accurate Food Image Segmentation
Swin-TUNA inserts layer-dependent depthwise-convolution adapters into a frozen Swin-L backbone and reports 50.56 mIoU on FoodSeg103 and 74.94 mIoU on UECFoodPix Complete with 8.13M trainable parameters.
Discussion (0). Continue with ORCID to comment.