Pith. sign in

REVIEW 2 cited by

Efficient World Models with Context-Aware Tokenization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.19320 v1 pith:TQ2FYC5X submitted 2024-06-27 cs.LG cs.AIcs.CV

classification cs.LGcs.AIcs.CV
keywords worldmodelsdeltadeltasirismodellingstatetokens
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Scaling up deep Reinforcement Learning (RL) methods presents a significant challenge. Following developments in generative modelling, model-based RL positions itself as a strong contender. Recent advances in sequence modelling have led to effective transformer-based world models, albeit at the price of heavy computations due to the long sequences of tokens required to accurately simulate environments. In this work, we propose $\Delta$-IRIS, a new agent with a world model architecture composed of a discrete autoencoder that encodes stochastic deltas between time steps and an autoregressive transformer that predicts future deltas by summarizing the current state of the world with continuous tokens. In the Crafter benchmark, $\Delta$-IRIS sets a new state of the art at multiple frame budgets, while being an order of magnitude faster to train than previous attention-based approaches. We release our code and models at https://github.com/vmicheli/delta-iris.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Can We Predict Your Next Move Without Breaking Your Privacy?

    cs.LG 2025-07 reject novelty 6.0 of 10

    A federated learning plus frozen LLM framework with outer-product aggregation claims top next-location prediction accuracy, but key reported numbers and an aggregation step undermine the claim.

  2. Improving Transformer World Models for Data-Efficient RL

    cs.LG 2025-02 conditional novelty 5.0 of 10

    A transformer world model agent using a static patch tokenizer, warmup before imagination training, and block teacher forcing reaches 69.66% reward on Craftax-classic, beating DreamerV3 and the human expert figure.

Pith tools