Pith. sign in

REVIEW 3 cited by

Grounding Multimodal Large Language Models in Actions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.07904 v2 pith:IUJGPJPP submitted 2024-06-12 cs.LG

classification cs.LG
keywords actionsactionmllmmultimodalspacebestdifferentembodied
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal Large Language Models (MLLMs) have demonstrated a wide range of capabilities across many domains, including Embodied AI. In this work, we study how to best ground a MLLM into different embodiments and their associated action spaces, with the goal of leveraging the multimodal world knowledge of the MLLM. We first generalize a number of methods through a unified architecture and the lens of action space adaptors. For continuous actions, we show that a learned tokenization allows for sufficient modeling precision, yielding the best performance on downstream tasks. For discrete actions, we demonstrate that semantically aligning these actions with the native output token space of the MLLM leads to the strongest performance. We arrive at these lessons via a thorough study of seven action space adapters on five different environments, encompassing over 114 embodied tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From Multimodal LLMs to Generalist Embodied Agents: Methods and Lessons

    cs.LG 2024-12 conditional novelty 6.0 of 10

    A single MLLM-based agent, finetuned with cross-domain supervision and online RL, achieves strong zero-shot generalization across manipulation, navigation, games, UI control, and planning.

  2. QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning

    cs.RO 2024-12 conditional novelty 5.0 of 10

    Compressing 10-step action chunks into discrete latent codes lets an 8B multimodal model drive a quadruped at controller frequency and raises average task success by about 65%.

  3. TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification

    cs.SD 2024-12 reject novelty 4.0 of 10

    Task-specific prompt ensembling with GPT-4-generated attributes and sources improves some zero-shot audio classification datasets while degrading others.

Pith tools