{"work":{"id":"6aeb260f-8c7c-4f9c-b98b-067cd7c59acd","openalex_id":"https://openalex.org/W4413116358","doi":"10.1126/sciadv.adu2488","arxiv_id":"2301.04104","raw_key":null,"title":"Mastering Diverse Domains through World Models","authors":null,"authors_text":"Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, Timothy Lillicrap","year":2023,"venue":"cs.AI","abstract":"Developing a general algorithm that learns to solve tasks across a wide range of applications has been a fundamental challenge in artificial intelligence. Although current reinforcement learning algorithms can be readily applied to tasks similar to what they have been developed for, configuring them for new application domains requires significant human expertise and experimentation. We present DreamerV3, a general algorithm that outperforms specialized methods across over 150 diverse tasks, with a single configuration. Dreamer learns a model of the environment and improves its behavior by imagining future scenarios. Robustness techniques based on normalization, balancing, and transformations enable stable learning across domains. Applied out of the box, Dreamer is the first algorithm to collect diamonds in Minecraft from scratch without human data or curricula. This achievement has been posed as a significant challenge in artificial intelligence that requires exploring farsighted strategies from pixels and sparse rewards in an open world. Our work allows solving challenging control problems without extensive experimentation, making reinforcement learning broadly applicable.","external_url":"https://arxiv.org/abs/2301.04104","cited_by_count":7,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2301.04104","created_at":"2026-05-09T22:54:16.628773+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Mastering Diverse Domains through World Models","render_title":"Mastering Diverse Domains through World Models"},"hub":{"state":{"work_id":"6aeb260f-8c7c-4f9c-b98b-067cd7c59acd","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":233,"external_cited_by_count":7,"distinct_field_count":15,"first_pith_cited_at":"2023-02-03T06:06:27+00:00","last_pith_cited_at":"2026-07-08T19:41:19+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T10:49:25.480484+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":31},{"context_role":"baseline","n":4},{"context_role":"method","n":4}],"polarity_counts":[{"context_polarity":"background","n":31},{"context_polarity":"baseline","n":4},{"context_polarity":"use_method","n":4}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Mastering Diverse Domains through World Models","claims":[{"claim_text":"Developing a general algorithm that learns to solve tasks across a wide range of applications has been a fundamental challenge in artificial intelligence. Although current reinforcement learning algorithms can be readily applied to tasks similar to what they have been developed for, configuring them for new application domains requires significant human expertise and experimentation. We present DreamerV3, a general algorithm that outperforms specialized methods across over 150 diverse tasks, with a single configuration. Dreamer learns a model of the environment and improves its behavior by ima","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"task objective, reward form, or local observation changes, as long as the envi- ronmental dynamics and delay-propagation characteristics remain similar. This supports cross-task transfer and rapid adaptation in random-delay environments. We systematically evaluate CausalDreamer on multiple continuous-control tasks from the DeepMind Control Suite (DMC) [8], and compares it with baselines including DreamerV3[4], CWM[9], and SAC[10]. The experimental results show that CausalDreamer generally outper","claim_type":"baseline","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"on reasoning-heavy tasks (GSM8K, MATH, HumanEval) despite ¡1% perplexity increase on WikiText-2. Reasoning tasks exhibit 2-3×larger degradation compared to knowledge-based tasks (MMLU), demonstrating that perplexity is insufficient for predicting downstream performance. Base model performance from official technical reports [149, 75]. Benchmarks: MMLU [62] (5-shot), GSM8K [27] (8-shot CoT), MATH [62] (4-shot), HumanEval [22] (0-shot pass@1). Model Task Type FP16 INT8 INT4 (GPTQ)∆INT4 LLaMA-2-70B","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"arXiv preprint arXiv:2010.02193 (2020). [20] Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. 2023. Mas- tering diverse domains through world models.arXiv preprint arXiv:2301.04104 (2023). [21] Nicklas Hansen, Hao Su, and Xiaolong Wang. 2023. Td-mpc2: Scalable, robust world models for continuous control.arXiv preprint arXiv:2310.16828(2023). [22] Nicklas Hansen, Xiaolong Wang, and Hao Su. 2022. Temporal difference learning for model predictive control.arXiv preprint arXiv:2203.","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"org/abs/1912. 01603. [16] Danijar Hafner, Kuang-Huei Lee, Ian Fischer, and Pieter Abbeel. Deep hierarchical planning from pixels.Advances in Neural Information Processing Systems, 35:26091-26104, 2022. [17] Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. Mastering diverse domains through world models.arXiv preprint arXiv:2301.04104, 2023. [18] Nicklas Hansen, Hao Su, and Xiaolong Wang. Td-mpc2: Scalable, robust world models for continuous control.arXiv preprint arXiv:2310.1682","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Brace: A benchmark for robust audio caption quality evaluation.arXiv preprint arXiv:2512.10403, 2025. [40] David Ha and Jürgen Schmidhuber. World models.arXiv preprint arXiv:1803.10122, 2(3), 2018. [41] Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. Dream to control: Learning behaviors by latent imagination.arXiv preprint arXiv:1912.01603, 2019. [42] Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. Mastering diverse domains through world models. arXiv prepri","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Three modules are required to implement this framework: a world model built on vari- ational free energy minimization, a self-prior that models the density of familiar sensory experiences, and a policy network built on expected free energy minimization (Fig. 2). The implementation builds on STORM [10], which uses a DreamerV3-like training pipeline [11] with transformer- based sequence modeling, and extends the core computations of active inference to high-dimensional problems via deep neural net","claim_type":"method","confidence":0.85,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Mastering Diverse Domains through World Models because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (27 contexts).","role_counts":[{"n":27,"context_role":"background"},{"n":4,"context_role":"method"},{"n":3,"context_role":"baseline"}]},"error":null,"updated_at":"2026-05-20T05:51:49.036477+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"76dee75f-541d-43e5-8504-725d910cf1b8","orcid":null,"display_name":"Danijar Hafner"},{"id":"7a5b0d5f-498b-41d1-8142-ec4a31dcb520","orcid":null,"display_name":"Jurgis Pasukonis"},{"id":"3651b710-1f6d-4de6-9dd2-f9e8628b782a","orcid":null,"display_name":"Jimmy Ba"},{"id":"610aa940-9c9c-4bdd-8cdb-ce3aa7519034","orcid":null,"display_name":"Timothy Lillicrap"}]},"error":null,"updated_at":"2026-05-20T05:51:49.028916+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T10:29:00.002639+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"World Models","work_id":"07227eee-8445-4c98-bce4-c6a6fd5ed907","shared_citers":26},{"title":"Dream to Control: Learning Behaviors by Latent Imagination","work_id":"5103f4be-344a-4139-8504-eaa59f5bac9d","shared_citers":17},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model","work_id":"3e7e65c5-5aed-4fe9-8414-2092bcb31cc7","shared_citers":13},{"title":"Ahmed Hendawy, Jan Peters, and Carlo D’Eramo","work_id":"360ec5fb-79fd-4490-bc73-3d161609c42d","shared_citers":12},{"title":"V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning","work_id":"a9c28401-f16a-4933-89f0-788e2f94e52b","shared_citers":12},{"title":"//arxiv.org/abs/2010.02193","work_id":"154f6f5f-bb34-456d-8107-45d5b51433ce","shared_citers":11},{"title":"$\\pi_0$: A Vision-Language-Action Flow Model for General Robot Control","work_id":"f790abdc-a796-482f-a40d-f8ee035ecfc2","shared_citers":10},{"title":"$\\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization","work_id":"d1ad7304-d09a-49bc-809e-846439f6aff9","shared_citers":10},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":9},{"title":"RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control","work_id":"ff438a8a-8003-4fae-9131-acd418b3597b","shared_citers":9},{"title":"RT-1: Robotics Transformer for Real-World Control at Scale","work_id":"e11bda85-8531-46bc-a07f-d0ade3643ab1","shared_citers":8},{"title":"//arxiv.org/abs/2310.06114","work_id":"16f38691-7ab6-4e23-bba5-6b656579e579","shared_citers":7},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":7},{"title":"Cosmos World Foundation Model Platform for Physical AI","work_id":"a2dba24c-318d-476a-8b21-4289c265810c","shared_citers":7},{"title":"DeepMind Control Suite","work_id":"54294ef0-c651-4d5a-a72b-f85a88329a71","shared_citers":7},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":7},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":6},{"title":"DINO-WM: World models on pre-trained vi- sual features enable zero-shot planning","work_id":"4a946586-a786-46da-9388-197c5410bf39","shared_citers":6},{"title":"GAIA-1: A Generative World Model for Autonomous Driving","work_id":"313484e6-a442-4522-8e19-d07e502844a8","shared_citers":6},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":6},{"title":"LeWorld- Model: Stable end-to-end joint-embedding predictive architecture from pixels","work_id":"d00a2f82-c871-4a58-9f34-22167e8efa93","shared_citers":6},{"title":"Octo: An Open-Source Generalist Robot Policy","work_id":"f9ca0722-8855-48c3-a27a-0eefb7e19253","shared_citers":6},{"title":"PaLM-E: An Embodied Multimodal Language Model","work_id":"5b99811a-1d93-47e2-9d59-f4045a0b74a2","shared_citers":6},{"title":"Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations","work_id":"62dbe235-8473-4190-8686-17e7437de50f","shared_citers":6}],"time_series":[{"n":1,"year":2023},{"n":1,"year":2024},{"n":1,"year":2025},{"n":59,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T10:39:24.493932+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T10:28:53.725038+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Mastering Diverse Domains through World Models","claims":[{"claim_text":"Developing a general algorithm that learns to solve tasks across a wide range of applications has been a fundamental challenge in artificial intelligence. Although current reinforcement learning algorithms can be readily applied to tasks similar to what they have been developed for, configuring them for new application domains requires significant human expertise and experimentation. We present DreamerV3, a general algorithm that outperforms specialized methods across over 150 diverse tasks, with a single configuration. Dreamer learns a model of the environment and improves its behavior by ima","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"task objective, reward form, or local observation changes, as long as the envi- ronmental dynamics and delay-propagation characteristics remain similar. This supports cross-task transfer and rapid adaptation in random-delay environments. We systematically evaluate CausalDreamer on multiple continuous-control tasks from the DeepMind Control Suite (DMC) [8], and compares it with baselines including DreamerV3[4], CWM[9], and SAC[10]. The experimental results show that CausalDreamer generally outper","claim_type":"baseline","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"on reasoning-heavy tasks (GSM8K, MATH, HumanEval) despite ¡1% perplexity increase on WikiText-2. Reasoning tasks exhibit 2-3×larger degradation compared to knowledge-based tasks (MMLU), demonstrating that perplexity is insufficient for predicting downstream performance. Base model performance from official technical reports [149, 75]. Benchmarks: MMLU [62] (5-shot), GSM8K [27] (8-shot CoT), MATH [62] (4-shot), HumanEval [22] (0-shot pass@1). Model Task Type FP16 INT8 INT4 (GPTQ)∆INT4 LLaMA-2-70B","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"arXiv preprint arXiv:2010.02193 (2020). [20] Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. 2023. Mas- tering diverse domains through world models.arXiv preprint arXiv:2301.04104 (2023). [21] Nicklas Hansen, Hao Su, and Xiaolong Wang. 2023. Td-mpc2: Scalable, robust world models for continuous control.arXiv preprint arXiv:2310.16828(2023). [22] Nicklas Hansen, Xiaolong Wang, and Hao Su. 2022. Temporal difference learning for model predictive control.arXiv preprint arXiv:2203.","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"org/abs/1912. 01603. [16] Danijar Hafner, Kuang-Huei Lee, Ian Fischer, and Pieter Abbeel. Deep hierarchical planning from pixels.Advances in Neural Information Processing Systems, 35:26091-26104, 2022. [17] Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. Mastering diverse domains through world models.arXiv preprint arXiv:2301.04104, 2023. [18] Nicklas Hansen, Hao Su, and Xiaolong Wang. Td-mpc2: Scalable, robust world models for continuous control.arXiv preprint arXiv:2310.1682","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Brace: A benchmark for robust audio caption quality evaluation.arXiv preprint arXiv:2512.10403, 2025. [40] David Ha and Jürgen Schmidhuber. World models.arXiv preprint arXiv:1803.10122, 2(3), 2018. [41] Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Mohammad Norouzi. Dream to control: Learning behaviors by latent imagination.arXiv preprint arXiv:1912.01603, 2019. [42] Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. Mastering diverse domains through world models. arXiv prepri","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Three modules are required to implement this framework: a world model built on vari- ational free energy minimization, a self-prior that models the density of familiar sensory experiences, and a policy network built on expected free energy minimization (Fig. 2). The implementation builds on STORM [10], which uses a DreamerV3-like training pipeline [11] with transformer- based sequence modeling, and extends the core computations of active inference to high-dimensional problems via deep neural net","claim_type":"method","confidence":0.85,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Mastering Diverse Domains through World Models because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (27 contexts).","role_counts":[{"n":27,"context_role":"background"},{"n":4,"context_role":"method"},{"n":3,"context_role":"baseline"}]},"error":null,"updated_at":"2026-05-20T05:51:49.032799+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Mastering Diverse Domains through World Models","claims":[{"claim_text":"Developing a general algorithm that learns to solve tasks across a wide range of applications has been a fundamental challenge in artificial intelligence. Although current reinforcement learning algorithms can be readily applied to tasks similar to what they have been developed for, configuring them for new application domains requires significant human expertise and experimentation. We present DreamerV3, a general algorithm that outperforms specialized methods across over 150 diverse tasks, with a single configuration. Dreamer learns a model of the environment and improves its behavior by ima","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Mastering Diverse Domains through World Models because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T10:39:24.495966+00:00"}},"summary":{"title":"Mastering Diverse Domains through World Models","claims":[{"claim_text":"Developing a general algorithm that learns to solve tasks across a wide range of applications has been a fundamental challenge in artificial intelligence. Although current reinforcement learning algorithms can be readily applied to tasks similar to what they have been developed for, configuring them for new application domains requires significant human expertise and experimentation. We present DreamerV3, a general algorithm that outperforms specialized methods across over 150 diverse tasks, with a single configuration. Dreamer learns a model of the environment and improves its behavior by ima","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Mastering Diverse Domains through World Models because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"World Models","work_id":"07227eee-8445-4c98-bce4-c6a6fd5ed907","shared_citers":26},{"title":"Dream to Control: Learning Behaviors by Latent Imagination","work_id":"5103f4be-344a-4139-8504-eaa59f5bac9d","shared_citers":17},{"title":"OpenVLA: An Open-Source Vision-Language-Action Model","work_id":"3e7e65c5-5aed-4fe9-8414-2092bcb31cc7","shared_citers":13},{"title":"Ahmed Hendawy, Jan Peters, and Carlo D’Eramo","work_id":"360ec5fb-79fd-4490-bc73-3d161609c42d","shared_citers":12},{"title":"V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning","work_id":"a9c28401-f16a-4933-89f0-788e2f94e52b","shared_citers":12},{"title":"//arxiv.org/abs/2010.02193","work_id":"154f6f5f-bb34-456d-8107-45d5b51433ce","shared_citers":11},{"title":"$\\pi_0$: A Vision-Language-Action Flow Model for General Robot Control","work_id":"f790abdc-a796-482f-a40d-f8ee035ecfc2","shared_citers":10},{"title":"$\\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization","work_id":"d1ad7304-d09a-49bc-809e-846439f6aff9","shared_citers":10},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":9},{"title":"RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control","work_id":"ff438a8a-8003-4fae-9131-acd418b3597b","shared_citers":9},{"title":"RT-1: Robotics Transformer for Real-World Control at Scale","work_id":"e11bda85-8531-46bc-a07f-d0ade3643ab1","shared_citers":8},{"title":"//arxiv.org/abs/2310.06114","work_id":"16f38691-7ab6-4e23-bba5-6b656579e579","shared_citers":7},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":7},{"title":"Cosmos World Foundation Model Platform for Physical AI","work_id":"a2dba24c-318d-476a-8b21-4289c265810c","shared_citers":7},{"title":"DeepMind Control Suite","work_id":"54294ef0-c651-4d5a-a72b-f85a88329a71","shared_citers":7},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":7},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":6},{"title":"DINO-WM: World models on pre-trained vi- sual features enable zero-shot planning","work_id":"4a946586-a786-46da-9388-197c5410bf39","shared_citers":6},{"title":"GAIA-1: A Generative World Model for Autonomous Driving","work_id":"313484e6-a442-4522-8e19-d07e502844a8","shared_citers":6},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":6},{"title":"LeWorld- Model: Stable end-to-end joint-embedding predictive architecture from pixels","work_id":"d00a2f82-c871-4a58-9f34-22167e8efa93","shared_citers":6},{"title":"Octo: An Open-Source Generalist Robot Policy","work_id":"f9ca0722-8855-48c3-a27a-0eefb7e19253","shared_citers":6},{"title":"PaLM-E: An Embodied Multimodal Language Model","work_id":"5b99811a-1d93-47e2-9d59-f4045a0b74a2","shared_citers":6},{"title":"Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations","work_id":"62dbe235-8473-4190-8686-17e7437de50f","shared_citers":6}],"time_series":[{"n":1,"year":2023},{"n":1,"year":2024},{"n":1,"year":2025},{"n":59,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"76dee75f-541d-43e5-8504-725d910cf1b8","orcid":null,"display_name":"Danijar Hafner","source":"manual","import_confidence":0.72},{"id":"3651b710-1f6d-4de6-9dd2-f9e8628b782a","orcid":null,"display_name":"Jimmy Ba","source":"manual","import_confidence":0.72},{"id":"7a5b0d5f-498b-41d1-8142-ec4a31dcb520","orcid":null,"display_name":"Jurgis Pasukonis","source":"manual","import_confidence":0.72},{"id":"610aa940-9c9c-4bdd-8cdb-ce3aa7519034","orcid":null,"display_name":"Timothy Lillicrap","source":"manual","import_confidence":0.72}]}}