{"work":{"id":"7529df29-9980-4a8a-b55e-307e3c2f357b","openalex_id":"https://openalex.org/W4298187450","doi":"10.48550/arxiv.2209.14988","arxiv_id":"2209.14988","raw_key":null,"title":"DreamFusion: Text-to-3D using 2D Diffusion","authors":null,"authors_text":"Ben Poole, Ajay Jain, Jonathan T. Barron, Ben Mildenhall","year":2022,"venue":"cs.CV","abstract":"Recent breakthroughs in text-to-image synthesis have been driven by diffusion models trained on billions of image-text pairs. Adapting this approach to 3D synthesis would require large-scale datasets of labeled 3D data and efficient architectures for denoising 3D data, neither of which currently exist. In this work, we circumvent these limitations by using a pretrained 2D text-to-image diffusion model to perform text-to-3D synthesis. We introduce a loss based on probability density distillation that enables the use of a 2D diffusion model as a prior for optimization of a parametric image generator. Using this loss in a DeepDream-like procedure, we optimize a randomly-initialized 3D model (a Neural Radiance Field, or NeRF) via gradient descent such that its 2D renderings from random angles achieve a low loss. The resulting 3D model of the given text can be viewed from any angle, relit by arbitrary illumination, or composited into any 3D environment. Our approach requires no 3D training data and no modifications to the image diffusion model, demonstrating the effectiveness of pretrained image diffusion models as priors.","external_url":"https://arxiv.org/abs/2209.14988","cited_by_count":467,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2209.14988","created_at":"2026-05-10T02:48:27.445779+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"DreamFusion: Text-to-3D using 2D Diffusion","render_title":"DreamFusion: Text-to-3D using 2D Diffusion"},"hub":{"state":{"work_id":"7529df29-9980-4a8a-b55e-307e3c2f357b","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":155,"external_cited_by_count":467,"distinct_field_count":9,"first_pith_cited_at":"2022-10-01T21:35:11+00:00","last_pith_cited_at":"2026-07-07T17:59:50+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T19:49:24.245360+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":17},{"context_role":"method","n":2},{"context_role":"dataset","n":1},{"context_role":"other","n":1}],"polarity_counts":[{"context_polarity":"background","n":16},{"context_polarity":"use_method","n":2},{"context_polarity":"support","n":1},{"context_polarity":"unclear","n":1},{"context_polarity":"use_dataset","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"DreamFusion: Text-to-3D using 2D Diffusion","claims":[{"claim_text":"Recent breakthroughs in text-to-image synthesis have been driven by diffusion models trained on billions of image-text pairs. Adapting this approach to 3D synthesis would require large-scale datasets of labeled 3D data and efficient architectures for denoising 3D data, neither of which currently exist. In this work, we circumvent these limitations by using a pretrained 2D text-to-image diffusion model to perform text-to-3D synthesis. We introduce a loss based on probability density distillation that enables the use of a 2D diffusion model as a prior for optimization of a parametric image gener","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"MLP-less/Explicit (III-G1) Plenoxels (2021) [43], DVGO (2021) [44], TensoRF (2022) [45] Sparse View (Sec. III-C) Cost V olume MVSNeRF [48] (2021), PixelNeRF (2020) [74], NeuRay (2021) [75] Others DietNeRF (2021) [49], DS-NeRF (2021) [23] Generative/ Conditional (Sec. III-D) GAN GIRAFFE (2020) [76], GRAF (2020) [77], π-GAN (2020) [78] , GNeRF (2021) [79], Stylenerf (2022) [80], EG3D (2022) [81] Diffusion DreamFusion (2022) [82], Magic3D (2022) [83], RealFusion (2023) [84] GLO NeRF-W (2020) [85], ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"ation of complex scenes with multiple objects and rich structural layouts [25-27]. We review these two lines of work below. 3D object generation. Owing to the absence of large-scale datasets with precise 3D annotations, early approaches to 3D object generation primarily relied on pre- trained 2D generative priors and optimization- based frameworks without explicit 3D supervi- sion [16, 17, 59-61]. For example, NeuralLift- 360 [60] optimizes a NeRF representation guided by CLIP, enforcing semanti","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"MS2Mesh-XR: Multi-modal Sketch-to-Mesh Generation in XR Environments. arXiv:2412.09008 [cs.CV] https://arxiv.org/abs/2412.09008 [22] Portia Wang, Mark R. Miller, Jeremy N. Bailenson, et al. 2024. Understanding virtual design behaviors: A large-scale analysis of the design process in Virtual Reality.Design Studies(2024). https://vhil.stanford.edu/sites/g/files/sbiybj29011/ files/media/file/design-studies-wang.pdf [23] Qiang Zou, Zhihong Tang, Hsi-Yung Feng, Shuming Gao, Chenchu Zhou, and Yusheng ","claim_type":"other","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"with transformers. InProceedings of the IEEE/CVF inter- national conference on computer vision, pages 4195-4205, 2023. 8 [36] Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M ¨uller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion mod- els for high-resolution image synthesis.arXiv preprint arXiv:2307.01952, 2023. 2 [37] Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Milden- hall. Dreamfusion: Text-to-3d using 2d diffusion.arXiv preprint arXiv:220","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"-Extensive experiments demonstrate that PAD not only achieves state-of- the-art performance in single-object generation and pose accuracy but also naturally enables high-quality compositional scene generation. 4 Zhou et al. 2 Related works 2.1 3D object generation Driven by the rapid advancement of powerful 2D image synthesis models [33,35], pioneering3Dgenerativeframeworks[3,23,34,41,49]haveextensivelyutilizedop- timization methods based on differentiable rendering [19,29,30] to lift 2D gener- ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Accessed 2025- 11-12. 5 [19] Nobuyuki Otsu. A threshold selection method from gray- level histograms.IEEE Transactions on Systems, Man, and Cybernetics, 9(1):62-66, 1979. 4 [20] William Peebles and Saining Xie. Scalable diffusion models with transformers. InProceedings of the IEEE/CVF inter- national conference on computer vision, pages 4195-4205, 2023. 2 [21] Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Milden- hall. Dreamfusion: Text-to-3d using 2d diffusion.arXiv preprint arXiv:2209.14988","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks DreamFusion: Text-to-3D using 2D Diffusion because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (17 contexts).","role_counts":[{"n":17,"context_role":"background"},{"n":2,"context_role":"method"},{"n":1,"context_role":"dataset"},{"n":1,"context_role":"other"}]},"error":null,"updated_at":"2026-06-29T16:29:09.234210+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"eb2faec5-1621-4088-bc26-f7fa3009e416","orcid":null,"display_name":"Ben Poole"},{"id":"95ac7a6c-c4de-4013-9786-7e5f1850c58e","orcid":null,"display_name":"Ajay Jain"},{"id":"9b070e08-fd1f-47be-983c-9b7811c7a5bb","orcid":null,"display_name":"Jonathan T. Barron"},{"id":"2427243a-45a5-4ecf-b7cb-588694e7fca2","orcid":null,"display_name":"Ben Mildenhall"}]},"error":null,"updated_at":"2026-06-29T16:29:09.623578+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T17:48:45.777441+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Dreamgaussian: Generative gaussian splatting for effi- cient 3d content creation.arXiv preprint arXiv:2309.16653","work_id":"790b039f-2f31-4657-93ac-f802a85cd72a","shared_citers":10},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":8},{"title":"SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis","work_id":"8034c587-fba6-4941-87ba-c98f2ac962cb","shared_citers":8},{"title":"Denoising Diffusion Implicit Models","work_id":"8fa2128b-d18c-405c-ac92-0e669cf89ac0","shared_citers":7},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":7},{"title":"InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models","work_id":"fc55eabb-0871-4dc4-8bab-b572e0d2aac4","shared_citers":7},{"title":"arXiv preprint arXiv:2506.16504 (2025)","work_id":"94fe94cc-10de-4093-be65-3170f7e638cb","shared_citers":6},{"title":"Lrm: Large reconstruction model for single image to 3d","work_id":"0662dc2c-cc1c-4358-99bd-2a5f34795738","shared_citers":6},{"title":"Mvdream: Multi-view diffusion for 3d gen- eration","work_id":"269d1cf4-9ad9-4b6b-8f5f-da1567956dd6","shared_citers":6},{"title":"SAM 3D: 3Dfy Anything in Images","work_id":"dc22e9ff-fcf5-4069-8ace-35ae3a0bfd7c","shared_citers":6},{"title":"Syncdreamer: Gen- erating multiview-consistent images from a single-view im- age","work_id":"b0c47120-5c87-4e79-8fe8-1be8e49f855f","shared_citers":6},{"title":"arXiv preprint arXiv:2212.08751(2022)","work_id":"9d7f0b29-b9ca-457f-9518-4a1506b43369","shared_citers":5},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":5},{"title":"Score-Based Generative Modeling through Stochastic Differential Equations","work_id":"d9110e53-a5d4-4794-a4c5-a575e91c31ad","shared_citers":5},{"title":"ShapeNet: An Information-Rich 3D Model Repository","work_id":"b2ac5b60-daa9-435b-9369-12271e126edd","shared_citers":5},{"title":"Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets","work_id":"4f68eada-27e3-437a-a2fe-6e4ca524d0d3","shared_citers":5},{"title":"arXiv2502.06608(2025) 5, 6, 10","work_id":"bf744acd-a6e4-4ba7-98eb-43811d95fa81","shared_citers":4},{"title":"Depth Anything 3: Recovering the Visual Space from Any Views","work_id":"0a54b500-1e9d-46c2-85eb-8e16cbac8461","shared_citers":4},{"title":"Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow","work_id":"a1989e1b-d66d-4533-be3a-fb9c5fd62290","shared_citers":4},{"title":"FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space","work_id":"5dfe19d5-3541-4803-8fe9-3c8b9e29b281","shared_citers":4},{"title":"Hi3dgen: High-fidelity 3d geometry generation from im- ages via normal bridging.arXiv preprint arXiv:2503.22236","work_id":"67aa4434-ef33-4b27-ae4f-1e5fa4225745","shared_citers":4},{"title":"Hierarchical Text-Conditional Image Generation with CLIP Latents","work_id":"0c6a768b-70b8-4242-bb0e-459f1008c9fc","shared_citers":4},{"title":"Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation","work_id":"c2c8fc56-8c16-45fa-a5d6-23c9c3d0720b","shared_citers":4},{"title":"Instant3d: Fast text-to-3d with sparse-view generation and large reconstruction model","work_id":"0c23e789-12f2-4cd9-900d-b96ddd3b811f","shared_citers":4}],"time_series":[{"n":39,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T17:49:14.199908+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T17:49:10.049015+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"DreamFusion: Text-to-3D using 2D Diffusion","claims":[{"claim_text":"Recent breakthroughs in text-to-image synthesis have been driven by diffusion models trained on billions of image-text pairs. Adapting this approach to 3D synthesis would require large-scale datasets of labeled 3D data and efficient architectures for denoising 3D data, neither of which currently exist. In this work, we circumvent these limitations by using a pretrained 2D text-to-image diffusion model to perform text-to-3D synthesis. We introduce a loss based on probability density distillation that enables the use of a 2D diffusion model as a prior for optimization of a parametric image gener","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"MLP-less/Explicit (III-G1) Plenoxels (2021) [43], DVGO (2021) [44], TensoRF (2022) [45] Sparse View (Sec. III-C) Cost V olume MVSNeRF [48] (2021), PixelNeRF (2020) [74], NeuRay (2021) [75] Others DietNeRF (2021) [49], DS-NeRF (2021) [23] Generative/ Conditional (Sec. III-D) GAN GIRAFFE (2020) [76], GRAF (2020) [77], π-GAN (2020) [78] , GNeRF (2021) [79], Stylenerf (2022) [80], EG3D (2022) [81] Diffusion DreamFusion (2022) [82], Magic3D (2022) [83], RealFusion (2023) [84] GLO NeRF-W (2020) [85], ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"ation of complex scenes with multiple objects and rich structural layouts [25-27]. We review these two lines of work below. 3D object generation. Owing to the absence of large-scale datasets with precise 3D annotations, early approaches to 3D object generation primarily relied on pre- trained 2D generative priors and optimization- based frameworks without explicit 3D supervi- sion [16, 17, 59-61]. For example, NeuralLift- 360 [60] optimizes a NeRF representation guided by CLIP, enforcing semanti","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"MS2Mesh-XR: Multi-modal Sketch-to-Mesh Generation in XR Environments. arXiv:2412.09008 [cs.CV] https://arxiv.org/abs/2412.09008 [22] Portia Wang, Mark R. Miller, Jeremy N. Bailenson, et al. 2024. Understanding virtual design behaviors: A large-scale analysis of the design process in Virtual Reality.Design Studies(2024). https://vhil.stanford.edu/sites/g/files/sbiybj29011/ files/media/file/design-studies-wang.pdf [23] Qiang Zou, Zhihong Tang, Hsi-Yung Feng, Shuming Gao, Chenchu Zhou, and Yusheng ","claim_type":"other","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"with transformers. InProceedings of the IEEE/CVF inter- national conference on computer vision, pages 4195-4205, 2023. 8 [36] Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas M ¨uller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion mod- els for high-resolution image synthesis.arXiv preprint arXiv:2307.01952, 2023. 2 [37] Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Milden- hall. Dreamfusion: Text-to-3d using 2d diffusion.arXiv preprint arXiv:220","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"-Extensive experiments demonstrate that PAD not only achieves state-of- the-art performance in single-object generation and pose accuracy but also naturally enables high-quality compositional scene generation. 4 Zhou et al. 2 Related works 2.1 3D object generation Driven by the rapid advancement of powerful 2D image synthesis models [33,35], pioneering3Dgenerativeframeworks[3,23,34,41,49]haveextensivelyutilizedop- timization methods based on differentiable rendering [19,29,30] to lift 2D gener- ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"Accessed 2025- 11-12. 5 [19] Nobuyuki Otsu. A threshold selection method from gray- level histograms.IEEE Transactions on Systems, Man, and Cybernetics, 9(1):62-66, 1979. 4 [20] William Peebles and Saining Xie. Scalable diffusion models with transformers. InProceedings of the IEEE/CVF inter- national conference on computer vision, pages 4195-4205, 2023. 2 [21] Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Milden- hall. Dreamfusion: Text-to-3d using 2d diffusion.arXiv preprint arXiv:2209.14988","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks DreamFusion: Text-to-3D using 2D Diffusion because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (17 contexts).","role_counts":[{"n":17,"context_role":"background"},{"n":2,"context_role":"method"},{"n":1,"context_role":"dataset"},{"n":1,"context_role":"other"}]},"error":null,"updated_at":"2026-06-29T16:29:09.625941+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"DreamFusion: Text-to-3D using 2D Diffusion","claims":[{"claim_text":"Recent breakthroughs in text-to-image synthesis have been driven by diffusion models trained on billions of image-text pairs. Adapting this approach to 3D synthesis would require large-scale datasets of labeled 3D data and efficient architectures for denoising 3D data, neither of which currently exist. In this work, we circumvent these limitations by using a pretrained 2D text-to-image diffusion model to perform text-to-3D synthesis. We introduce a loss based on probability density distillation that enables the use of a 2D diffusion model as a prior for optimization of a parametric image gener","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks DreamFusion: Text-to-3D using 2D Diffusion because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T17:49:10.053019+00:00"}},"summary":{"title":"DreamFusion: Text-to-3D using 2D Diffusion","claims":[{"claim_text":"Recent breakthroughs in text-to-image synthesis have been driven by diffusion models trained on billions of image-text pairs. Adapting this approach to 3D synthesis would require large-scale datasets of labeled 3D data and efficient architectures for denoising 3D data, neither of which currently exist. In this work, we circumvent these limitations by using a pretrained 2D text-to-image diffusion model to perform text-to-3D synthesis. We introduce a loss based on probability density distillation that enables the use of a 2D diffusion model as a prior for optimization of a parametric image gener","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks DreamFusion: Text-to-3D using 2D Diffusion because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Dreamgaussian: Generative gaussian splatting for effi- cient 3d content creation.arXiv preprint arXiv:2309.16653","work_id":"790b039f-2f31-4657-93ac-f802a85cd72a","shared_citers":10},{"title":"DINOv2: Learning Robust Visual Features without Supervision","work_id":"26b304e5-b54a-4f26-be7e-83299eca52e4","shared_citers":8},{"title":"SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis","work_id":"8034c587-fba6-4941-87ba-c98f2ac962cb","shared_citers":8},{"title":"Denoising Diffusion Implicit Models","work_id":"8fa2128b-d18c-405c-ac92-0e669cf89ac0","shared_citers":7},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":7},{"title":"InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models","work_id":"fc55eabb-0871-4dc4-8bab-b572e0d2aac4","shared_citers":7},{"title":"arXiv preprint arXiv:2506.16504 (2025)","work_id":"94fe94cc-10de-4093-be65-3170f7e638cb","shared_citers":6},{"title":"Lrm: Large reconstruction model for single image to 3d","work_id":"0662dc2c-cc1c-4358-99bd-2a5f34795738","shared_citers":6},{"title":"Mvdream: Multi-view diffusion for 3d gen- eration","work_id":"269d1cf4-9ad9-4b6b-8f5f-da1567956dd6","shared_citers":6},{"title":"SAM 3D: 3Dfy Anything in Images","work_id":"dc22e9ff-fcf5-4069-8ace-35ae3a0bfd7c","shared_citers":6},{"title":"Syncdreamer: Gen- erating multiview-consistent images from a single-view im- age","work_id":"b0c47120-5c87-4e79-8fe8-1be8e49f855f","shared_citers":6},{"title":"arXiv preprint arXiv:2212.08751(2022)","work_id":"9d7f0b29-b9ca-457f-9518-4a1506b43369","shared_citers":5},{"title":"Auto-Encoding Variational Bayes","work_id":"97d95295-30e1-42b4-bbf6-85f0fa4edb44","shared_citers":5},{"title":"Score-Based Generative Modeling through Stochastic Differential Equations","work_id":"d9110e53-a5d4-4794-a4c5-a575e91c31ad","shared_citers":5},{"title":"ShapeNet: An Information-Rich 3D Model Repository","work_id":"b2ac5b60-daa9-435b-9369-12271e126edd","shared_citers":5},{"title":"Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets","work_id":"4f68eada-27e3-437a-a2fe-6e4ca524d0d3","shared_citers":5},{"title":"arXiv2502.06608(2025) 5, 6, 10","work_id":"bf744acd-a6e4-4ba7-98eb-43811d95fa81","shared_citers":4},{"title":"Depth Anything 3: Recovering the Visual Space from Any Views","work_id":"0a54b500-1e9d-46c2-85eb-8e16cbac8461","shared_citers":4},{"title":"Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow","work_id":"a1989e1b-d66d-4533-be3a-fb9c5fd62290","shared_citers":4},{"title":"FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space","work_id":"5dfe19d5-3541-4803-8fe9-3c8b9e29b281","shared_citers":4},{"title":"Hi3dgen: High-fidelity 3d geometry generation from im- ages via normal bridging.arXiv preprint arXiv:2503.22236","work_id":"67aa4434-ef33-4b27-ae4f-1e5fa4225745","shared_citers":4},{"title":"Hierarchical Text-Conditional Image Generation with CLIP Latents","work_id":"0c6a768b-70b8-4242-bb0e-459f1008c9fc","shared_citers":4},{"title":"Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation","work_id":"c2c8fc56-8c16-45fa-a5d6-23c9c3d0720b","shared_citers":4},{"title":"Instant3d: Fast text-to-3d with sparse-view generation and large reconstruction model","work_id":"0c23e789-12f2-4cd9-900d-b96ddd3b811f","shared_citers":4}],"time_series":[{"n":39,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"95ac7a6c-c4de-4013-9786-7e5f1850c58e","orcid":null,"display_name":"Ajay Jain","source":"manual","import_confidence":0.72},{"id":"2427243a-45a5-4ecf-b7cb-588694e7fca2","orcid":null,"display_name":"Ben Mildenhall","source":"manual","import_confidence":0.72},{"id":"eb2faec5-1621-4088-bc26-f7fa3009e416","orcid":null,"display_name":"Ben Poole","source":"manual","import_confidence":0.72},{"id":"9b070e08-fd1f-47be-983c-9b7811c7a5bb","orcid":null,"display_name":"Jonathan T. Barron","source":"manual","import_confidence":0.72}]}}