REVIEW 10 cited by
COMBO: Compositional World Models for Embodied Multi-Agent Cooperation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
In this paper, we investigate the problem of embodied multi-agent cooperation, where decentralized agents must cooperate given only egocentric views of the world. To effectively plan in this setting, in contrast to learning world dynamics in a single-agent scenario, we must simulate world dynamics conditioned on an arbitrary number of agents' actions given only partial egocentric visual observations of the world. To address this issue of partial observability, we first train generative models to estimate the overall world state given partial egocentric observations. To enable accurate simulation of multiple sets of actions on this world state, we then propose to learn a compositional world model for multi-agent cooperation by factorizing the naturally composable joint actions of multiple agents and compositionally generating the video conditioned on the world state. By leveraging this compositional world model, in combination with Vision Language Models to infer the actions of other agents, we can use a tree search procedure to integrate these modules and facilitate online cooperative planning. We evaluate our methods on three challenging benchmarks with 2-4 agents. The results show our compositional world model is effective and the framework enables the embodied agents to cooperate efficiently with different agents across various tasks and an arbitrary number of agents, showing the promising future of our proposed methods. More videos can be found at https://umass-embodied-agi.github.io/COMBO/.
Forward citations
Cited by 10 Pith papers
-
TeamCraft: A Benchmark for Multi-Modal Multi-Agent Systems in Minecraft
TeamCraft presents a large multi-modal, multi-agent Minecraft benchmark and shows that current models generalize poorly to novel goals, scenes, and team sizes.
-
MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation
Credit assignment via LMM pairwise comparisons plus Bradley–Terry rank aggregation and potential-based shaping improves cooperative MARL under sparse rewards and dynamic agent counts.
-
VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
VLM4D benchmarks spatiotemporal reasoning in VLMs and finds large gaps versus humans, with proposed methods showing partial improvement.
-
"Harmless to You, Hurtful to Me!": Investigating the Detection of Toxic Languages Grounded in the Perspective of Youth
The authors construct the first Chinese youth-toxicity dataset, show that youth and adult perceptions of toxic language diverge, and report that adding contextual meta information improves detection accuracy.
-
TesserAct: Learning 4D Embodied World Models
A 4D embodied world model that generates RGB-depth-normal videos from an image and instruction, reconstructs the scene as point clouds, and uses those point clouds to train better robot manipulation policies.
-
Ego-centric Learning of Communicative World Models for Autonomous Driving
Sharing compressed latent states and planned waypoints between agents, triggered by prediction errors, improves multi-agent driving performance in CARLA while cutting communication bandwidth by roughly 50x.
-
Generative AI in Embodied Systems: System-Level Analysis of Performance, Efficiency and Scalability
A workload characterization of 14 embodied LLM agent systems showing that planning and communication dominate latency, communication is often redundant, and multi-agent systems scale poorly.
-
Effect of Adaptive Communication Support on LLM-powered Human-Robot Collaboration
A human-robot teaming framework with adjustable LLM feedback improves collaboration in easy and medium tasks, but overly frequent feedback from a less capable LLM hurts performance in hard tasks.
-
A Survey of Interactive Generative Video
A survey that divides interactive generative video research into five modules: generation, control, memory, dynamics, and intelligence.
-
GenAI-based Multi-Agent Reinforcement Learning towards Distributed Agent Intelligence: A Generative-RL Agent Perspective
A position paper claiming that generative-AI agents that model and predict multi-agent dynamics will replace today's reactive MARL approaches.
Discussion (0). Continue with ORCID to comment.