{"paper":{"title":"D4RL: Datasets for Deep Data-Driven Reinforcement Learning","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"New benchmark datasets for offline RL, drawn from human demonstrations and mixed policies, expose deficiencies in existing algorithms.","cross_cats":["stat.ML"],"primary_cat":"cs.LG","authors_text":"Aviral Kumar, George Tucker, Justin Fu, Ofir Nachum, Sergey Levine","submitted_at":"2020-04-15T17:18:19Z","abstract_excerpt":"The offline reinforcement learning (RL) setting (also known as full batch RL), where a policy is learned from a static dataset, is compelling as progress enables RL methods to take advantage of large, previously-collected datasets, much like how the rise of large datasets has fueled results in supervised learning. However, existing online RL benchmarks are not tailored towards the offline setting and existing offline RL benchmarks are restricted to data generated by partially-trained agents, making progress in offline RL difficult to measure. In this work, we introduce benchmarks specifically "},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"By moving beyond simple benchmark tasks and data collected by partially-trained RL agents, we reveal important and unappreciated deficiencies of existing algorithms.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That datasets generated via hand-designed controllers, human demonstrators, multitask settings, and mixtures of policies capture the key properties most relevant to real-world offline RL applications.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"D4RL supplies new offline RL benchmarks and datasets from expert and mixed sources to expose weaknesses in existing algorithms and standardize evaluation.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"New benchmark datasets for offline RL, drawn from human demonstrations and mixed policies, expose deficiencies in existing algorithms.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"0c06905e9ba24f8729ba91ab0b61bdb9ea59af869a8a84e57b82b8486d7ef737"},"source":{"id":"2004.07219","kind":"arxiv","version":4},"verdict":{"id":"f9c60159-366b-4c1b-bde8-534bda9e44c0","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-12T23:14:59.570931Z","strongest_claim":"By moving beyond simple benchmark tasks and data collected by partially-trained RL agents, we reveal important and unappreciated deficiencies of existing algorithms.","one_line_summary":"D4RL supplies new offline RL benchmarks and datasets from expert and mixed sources to expose weaknesses in existing algorithms and standardize evaluation.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That datasets generated via hand-designed controllers, human demonstrators, multitask settings, and mixtures of policies capture the key properties most relevant to real-world offline RL applications.","pith_extraction_headline":"New benchmark datasets for offline RL, drawn from human demonstrations and mixed policies, expose deficiencies in existing algorithms."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2004.07219/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":24,"sample":[{"doi":"","year":1908,"title":"Preprint arXiv:1908.00261 , year=","work_id":"f49e0fd8-840b-4c3d-bf03-e057e1da8f2d","ref_index":1,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":1909,"title":"Scaling data-driven robotics with reward sketching and batch reinforcement learning.Preprint arXiv:1909.12200","work_id":"c44fff83-b8c4-4bf1-a47d-0f5c0046b36c","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2018,"title":"End- to-end driving via conditional imitation learning","work_id":"a28935f5-b348-47b2-9aeb-25c303b9a4aa","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":1904,"title":"Challenges of Real-World Reinforcement Learning","work_id":"fc99449a-80f4-4f37-a028-7b3774c78bf6","ref_index":4,"cited_arxiv_id":"1904.12901","is_internal_anchor":false},{"doi":"","year":2003,"title":"Mankowitz, Jerry Li, Cosmin Paduraru, Sven Gowal, and Todd Hes- ter","work_id":"72539193-de5d-4093-a85f-abb4eba8699e","ref_index":5,"cited_arxiv_id":"","is_internal_anchor":false}],"resolved_work":24,"snapshot_sha256":"680ebf13a40fb439d9b1b23727eb2e41ac080d99d079f74ccf5b59c7c056f646","internal_anchors":7},"formal_canon":{"evidence_count":1,"snapshot_sha256":"896b842cd86dfef1170675a52ee16c796e1c4a6ec93fe924fa49b97263993227"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}