Pith. sign in

REVIEW 1 cited by

YOLO-MARL: You Only LLM Once for Multi-Agent Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.03997 v2 pith:BK35VD2F submitted 2024-10-05 cs.MA

classification cs.MA
keywords llmsmarlyolo-marlagentscooperativelearningonlycapabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Advancements in deep multi-agent reinforcement learning (MARL) have positioned it as a promising approach for decision-making in cooperative games. However, it still remains challenging for MARL agents to learn cooperative strategies for some game environments. Recently, large language models (LLMs) have demonstrated emergent reasoning capabilities, making them promising candidates for enhancing coordination among the agents. However, due to the model size of LLMs, it can be expensive to frequently infer LLMs for actions that agents can take. In this work, we propose You Only LLM Once for MARL (YOLO-MARL), a novel framework that leverages the high-level task planning capabilities of LLMs to improve the policy learning process of multi-agents in cooperative games. Notably, for each game environment, YOLO-MARL only requires one time interaction with LLMs in the proposed strategy generation, state interpretation and planning function generation modules, before the MARL policy training process. This avoids the ongoing costs and computational time associated with frequent LLMs API calls during training. Moreover, trained decentralized policies based on normal-sized neural networks operate independently of the LLM. We evaluate our method across two different environments and demonstrate that YOLO-MARL outperforms traditional MARL algorithms.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RALLY: Role-Adaptive LLM-Driven Yoked Navigation for Agentic UAV Swarms

    cs.MA 2025-07 conditional novelty 4.0 of 10

    RALLY couples a two-stage LLM consensus module with a QMIX-style role-assignment network and reports higher reward and better generalization than three baselines in drone-swarm coverage simulations.

Pith tools