Pith. sign in

REVIEW 6 cited by

Learning to Model the World with Language

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.01399 v2 pith:7TH7DWH4 submitted 2023-07-31 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords languageworldagentslearningmodeldiversedynalangfuture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To interact with humans and act in the world, agents need to understand the range of language that people use and relate it to the visual world. While current agents can learn to execute simple language instructions, we aim to build agents that leverage diverse language -- language like "this button turns on the TV" or "I put the bowls away" -- that conveys general knowledge, describes the state of the world, provides interactive feedback, and more. Our key idea is that agents should interpret such diverse language as a signal that helps them predict the future: what they will observe, how the world will behave, and which situations will be rewarded. This perspective unifies language understanding with future prediction as a powerful self-supervised learning objective. We instantiate this in Dynalang, an agent that learns a multimodal world model to predict future text and image representations, and learns to act from imagined model rollouts. While current methods that learn language-conditioned policies degrade in performance with more diverse types of language, we show that Dynalang learns to leverage environment descriptions, game rules, and instructions to excel on tasks ranging from game-playing to navigating photorealistic home scans. Finally, we show that our method enables additional capabilities due to learning a generative model: Dynalang can be pretrained on text-only data, enabling learning from offline datasets, and generate language grounded in an environment.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. YoCausal: How Far is Video Generation from World Model? A Causality Perspective

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    YoCausal benchmark shows video diffusion models detect the arrow of time but lack genuine causal understanding relative to humans.

  2. LabRobFail: A Benchmark for Robotic Failure Analysis in Chemical Self-driving Laboratory

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A chemical-lab failure sim, 20K-trajectory dataset, six-axis benchmark, and specialized VLM raise failure detection to 90.8% on seen scenes and lift downstream policy success by 4–16 points.

  3. Audio-Visual World Models: Learning Physically Grounded Multisensory Dynamics

    cs.MM 2025-11 conditional novelty 6.0 of 10

    Future video frames, binaural audio, and task rewards can be generated jointly under action control by a diffusion transformer trained on a new 30-hour simulated audio-visual benchmark, with small downstream navigation gains.

  4. GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation

    cs.RO 2025-06 unverdicted novelty 6.0 of 10

    GAF creates 4D dynamic scene models by adding motion to 3D Gaussians, enabling better reconstruction and 7.3% higher success in robotic tasks.

  5. Evaluating LLMs as Interpretable Controllers for Dynamical Systems

    cs.AI 2026-06 conditional novelty 5.0 of 10

    In a simulated thermal enclosure, high-capability LLMs achieve accurate setpoint tracking and coherent reasoning, while small models fail; giving the LLM the exact plant simulator as a prediction tool improves smoothn...

  6. LaGO: Latent Action Guidance for Online Reinforcement Learning

    cs.AI 2026-06 unverdicted novelty 4.0 of 10

    LaGO improves online RL success rates over vanilla PPO by using pretrained LLMs as latent action priors, raising rates from 15.1% to 27.2% on CLEVR-Robot and 2.7% to 15.2% on Meta-World.

Pith tools