Pith. sign in

REVIEW 2 cited by

FoodLMM: A Versatile Food Assistant using Large Multi-modal Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.14991 v2 pith:IR4C7FJI submitted 2023-12-22 cs.CV

classification cs.CV
keywords foodfoodlmmsegmentationlmmsmodelassistantbenchmarksconversation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Multi-modal Models (LMMs) have made impressive progress in many vision-language tasks. Nevertheless, the performance of general LMMs in specific domains is still far from satisfactory. This paper proposes FoodLMM, a versatile food assistant based on LMMs with various capabilities, including food recognition, ingredient recognition, recipe generation, nutrition estimation, food segmentation and multi-round conversation. To facilitate FoodLMM to deal with tasks beyond pure text output, we introduce a series of novel task-specific tokens and heads, enabling the model to predict food nutritional values and multiple segmentation masks. We adopt a two-stage training strategy. In the first stage, we utilize multiple public food benchmarks for multi-task learning by leveraging the instruct-following paradigm. In the second stage, we construct a multi-round conversation dataset and a reasoning segmentation dataset to fine-tune the model, enabling it to conduct professional dialogues and generate segmentation masks based on complex reasoning in the food domain. Our fine-tuned FoodLMM achieves state-of-the-art results across several food benchmarks. We will make our code, models and datasets publicly available.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SFOOD: A Multimodal Benchmark for Comprehensive Food Attribute Analysis Beyond RGB with Spectral Insights

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SFOOD combines existing food datasets with self-collected hyperspectral images to create a six-task benchmark, and its evaluations suggest spectral bands improve sweetness and herbal classification while current model...

  2. RecipeGen: A Step-Aligned Multimodal Benchmark for Real-World Recipe Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    RecipeGen is a new benchmark with 26,453 recipes, 196,724 step-aligned images, and 4,491 cooking videos, plus three domain-specific evaluation metrics for recipe generation.

Pith tools