An audio-conditioned STEVE-1 agent, built with a new Minecraft audio-video CLIP model and a learned prior, matches or beats text- and video-conditioned versions on most short-horizon collection tasks.
Behavioral Cloning via Search in Video PreTraining Latent Space
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Our aim is to build autonomous agents that can solve tasks in environments like Minecraft. To do so, we used an imitation learning-based approach. We formulate our control problem as a search problem over a dataset of experts' demonstrations, where the agent copies actions from a similar demonstration trajectory of image-action pairs. We perform a proximity search over the BASALT MineRL-dataset in the latent representation of a Video PreTraining model. The agent copies the actions from the expert trajectory as long as the distance between the state representations of the agent and the selected expert trajectory from the dataset do not diverge. Then the proximity search is repeated. Our approach can effectively recover meaningful demonstration trajectories and show human-like behavior of an agent in the Minecraft environment.
citation-role summary
citation-polarity summary
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
STEVE-Audio: Expanding the Goal Conditioning Modalities of Embodied Agents in Minecraft
An audio-conditioned STEVE-1 agent, built with a new Minecraft audio-video CLIP model and a learned prior, matches or beats text- and video-conditioned versions on most short-horizon collection tasks.