Pith. sign in

REVIEW 4 cited by

RoDE: Linear Rectified Mixture of Diverse Experts for Food Large Multi-Modal Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.12730 v3 pith:CWSVBDIG submitted 2024-07-17 cs.CV cs.AI

classification cs.CVcs.AI
keywords foodrodelmmstasksapproachdatadiverseexperts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Multi-modal Models (LMMs) have significantly advanced a variety of vision-language tasks. The scalability and availability of high-quality training data play a pivotal role in the success of LMMs. In the realm of food, while comprehensive food datasets such as Recipe1M offer an abundance of ingredient and recipe information, they often fall short of providing ample data for nutritional analysis. The Recipe1M+ dataset, despite offering a subset for nutritional evaluation, is limited in the scale and accuracy of nutrition information. To bridge this gap, we introduce Uni-Food, a unified food dataset that comprises over 100,000 images with various food labels, including categories, ingredients, recipes, and ingredient-level nutritional information. Uni-Food is designed to provide a more holistic approach to food data analysis, thereby enhancing the performance and capabilities of LMMs in this domain. To mitigate the conflicts arising from multi-task supervision during fine-tuning of LMMs, we introduce a novel Linear Rectification Mixture of Diverse Experts (RoDE) approach. RoDE utilizes a diverse array of experts to address tasks of varying complexity, thereby facilitating the coordination of trainable parameters, i.e., it allocates more parameters for more complex tasks and, conversely, fewer parameters for simpler tasks. RoDE implements linear rectification union to refine the router's functionality, thereby enhancing the efficiency of sparse task allocation. These design choices endow RoDE with features that ensure GPU memory efficiency and ease of optimization. Our experimental results validate the effectiveness of our proposed approach in addressing the inherent challenges of food-related multitasking.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. OmniFood8K: Single-Image Nutrition Estimation via Hierarchical Frequency-Aligned Fusion

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    OmniFood8K supplies a large Chinese-food nutrition dataset and a single-image model that predicts depth then hierarchically fuses RGB and depth features in frequency space for improved nutrition estimates.

  2. SFOOD: A Multimodal Benchmark for Comprehensive Food Attribute Analysis Beyond RGB with Spectral Insights

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SFOOD combines existing food datasets with self-collected hyperspectral images to create a six-task benchmark, and its evaluations suggest spectral bands improve sweetness and herbal classification while current model...

  3. CogniRoute: Learning to Route Social Evidence in Omni-Modal Models

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    CogniRoute adds a cognitive schema and route-aware RL to an omni-modal MoE, reaching 59.38% accuracy on a new 118K-example social video QA benchmark and beating prior baselines by 15-27 points.

  4. Food-R1: A Unified Multi-Task Food Vision-Language Model with Reinforcement Learning

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    Introduces CalorieBench-80K benchmark with CoT calorie reasoning and Food-R1 VLM trained via CoT cold-start then GRPO reinforcement fine-tuning, claiming consistent outperformance on food tasks.

Pith tools