{"paper":{"title":"The Replica Dataset: A Digital Replica of Indoor Spaces","license":"http://creativecommons.org/licenses/by/4.0/","headline":"Replica is a dataset of 18 photo-realistic 3D indoor scenes designed so machine learning models trained on it may work directly on real-world data.","cross_cats":["cs.GR","eess.IV"],"primary_cat":"cs.CV","authors_text":"Anton Clarkson, Brian Budge, Carl Ren, Dhruv Batra, Elias Mueggler, Erik Wijmans, Hauke M. Strasdat, Jakob J. Engel, Jesus Briales, Julian Straub, June Yon, Kimberly Leon, Lingni Ma, Luis Pesqueira, Manolis Savva, Michael Goesele, Mingfei Yan, Nigel Carter, Raul Mur-Artal, Renzo De Nardi, Richard Newcombe, Shobhit Verma, Simon Green, Steven Lovegrove, Thomas Whelan, Tyler Gillingham, Xiaqing Pan, Yajie Yan, Yufan Chen, Yuyang Zou","submitted_at":"2019-06-13T16:29:58Z","abstract_excerpt":"We introduce Replica, a dataset of 18 highly photo-realistic 3D indoor scene reconstructions at room and building scale. Each scene consists of a dense mesh, high-resolution high-dynamic-range (HDR) textures, per-primitive semantic class and instance information, and planar mirror and glass reflectors. The goal of Replica is to enable machine learning (ML) research that relies on visually, geometrically, and semantically realistic generative models of the world - for instance, egocentric computer vision, semantic segmentation in 2D and 3D, geometric inference, and the development of embodied a"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Due to the high level of realism of the renderings from Replica, there is hope that ML systems trained on Replica may transfer directly to real world image and video data.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the 18 scenes achieve sufficient photo-realism and geometric accuracy in their meshes, textures, and semantics for ML models to transfer directly to real-world data without additional domain adaptation.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Replica is a new dataset of 18 highly detailed 3D reconstructions of indoor spaces with meshes, high-resolution HDR textures, per-primitive semantics, and mirror/glass reflectors for realistic ML training.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Replica is a dataset of 18 photo-realistic 3D indoor scenes designed so machine learning models trained on it may work directly on real-world data.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"2c07e948db0c17e66edf4897af20ed9a0de50dc00f0821f0de0a276466629673"},"source":{"id":"1906.05797","kind":"arxiv","version":1},"verdict":{"id":"ca8c9a91-e69e-4105-b059-5d3a6bb33834","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-12T16:29:51.021545Z","strongest_claim":"Due to the high level of realism of the renderings from Replica, there is hope that ML systems trained on Replica may transfer directly to real world image and video data.","one_line_summary":"Replica is a new dataset of 18 highly detailed 3D reconstructions of indoor spaces with meshes, high-resolution HDR textures, per-primitive semantics, and mirror/glass reflectors for realistic ML training.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the 18 scenes achieve sufficient photo-realism and geometric accuracy in their meshes, textures, and semantics for ML models to transfer directly to real-world data without additional domain adaptation.","pith_extraction_headline":"Replica is a dataset of 18 photo-realistic 3D indoor scenes designed so machine learning models trained on it may work directly on real-world data."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/1906.05797/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":28,"sample":[{"doi":"","year":2018,"title":"On Evaluation of Embodied Navigation Agents","work_id":"3b074aa9-2ff9-4ad6-8796-6a25689ecfd3","ref_index":1,"cited_arxiv_id":"1807.06757","is_internal_anchor":true},{"doi":"","year":2018,"title":"Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments","work_id":"c286bd3d-a338-46c4-8109-0c4f9c542cfa","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2015,"title":"Lawrence Zitnick, and Devi Parikh","work_id":"535137bf-c032-453e-8cc8-7a175a62b510","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2016,"title":"Zamir, Helen Jiang, Ioannis Brilakis, Martin Fischer, and Silvio Savarese","work_id":"43a7813f-2984-43f4-b9a8-85fd72c01742","ref_index":4,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2008,"title":"Ptex: Per-face texture mapping for production rendering","work_id":"aa0cbca4-5e88-481d-8060-c83e53587d7e","ref_index":5,"cited_arxiv_id":"","is_internal_anchor":false}],"resolved_work":28,"snapshot_sha256":"172217c711b3477f838d372de4eecfcd09f00e02fec17cb999569cccbaf51842","internal_anchors":2},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}