REVIEW 20 cited by
ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Score distillation sampling (SDS) has shown great promise in text-to-3D generation by distilling pretrained large-scale text-to-image diffusion models, but suffers from over-saturation, over-smoothing, and low-diversity problems. In this work, we propose to model the 3D parameter as a random variable instead of a constant as in SDS and present variational score distillation (VSD), a principled particle-based variational framework to explain and address the aforementioned issues in text-to-3D generation. We show that SDS is a special case of VSD and leads to poor samples with both small and large CFG weights. In comparison, VSD works well with various CFG weights as ancestral sampling from diffusion models and simultaneously improves the diversity and sample quality with a common CFG weight (i.e., $7.5$). We further present various improvements in the design space for text-to-3D such as distillation time schedule and density initialization, which are orthogonal to the distillation algorithm yet not well explored. Our overall approach, dubbed ProlificDreamer, can generate high rendering resolution (i.e., $512\times512$) and high-fidelity NeRF with rich structure and complex effects (e.g., smoke and drops). Further, initialized from NeRF, meshes fine-tuned by VSD are meticulously detailed and photo-realistic. Project page and codes: https://ml.cs.tsinghua.edu.cn/prolificdreamer/
Forward citations
Cited by 20 Pith papers
-
D2PO: Optimizing Diffusion Samplers via Dynamic Preference
D2PO learns better low-NFE diffusion timestep schedules and CFG weights via DPO on a score-based energy with a dynamic denser-schedule preference target.
-
MVGBench: Comprehensive Benchmark for Multi-view Generation Models
MVGBench evaluates multi-view generators through self-consistency of 3D reconstructions and uses this protocol to rank 12 models and build a better one.
-
GPS as a Control Signal for Image Generation
A diffusion model conditioned on GPS tags and text can generate location-specific images and reconstruct 3D landmarks via score distillation sampling, without explicit pose estimation.
-
SynCity 3000: Bootstrapping Scene-Scale 3D Diffusion
SynCity 3000 generates large, coherent 3D scenes from text by fine-tuning an image-to-3D diffusion model to operate convolutionally on overlapping windows, trained on procedurally generated synthetic scene data.
-
PixGS: Pixel-Space Diffusion for Direct 3D Gaussian Splat Generation
PixGS is a single-stage pixel-space diffusion model that directly produces high-quality 3D Gaussian Splats from text or images in ~1s, outperforming multi-stage latent methods on standard benchmarks.
-
Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation
A frozen VLM's dual-query Yes/No log-odds act as a differentiable semantic-and-spatial critic, improving alignment and geometry in both SDS-based and feed-forward text-to-3D pipelines.
-
Discrete Noise Inversion for Next-scale Autoregressive Text-based Image Editing
VARIN uses a Location-aware Argmax Inversion pseudo-inverse of Gumbel-max sampling to extract editable discrete noises, enabling training-free prompt-guided editing for visual autoregressive models.
-
Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis
A video-to-4D model that encodes mesh animations into compact Gaussian variation latents and diffuses them conditioned on the video and a canonical Gaussian splat.
-
AutoPartGen: Autogressive 3D Part Generation and Discovery
AutoPartGen generates 3D objects as a sequence of latent-space parts, conditioning each new part on previously generated parts, and reports state-of-the-art part completion on PartObjaverse-Tiny.
-
SeqTex: Generate Mesh Textures in Video Sequence
SeqTex adapts a pretrained video diffusion model to directly generate complete UV texture maps by jointly predicting four multi-view images and the UV map as a five-frame sequence.
-
Matrix3D: Large Photogrammetry Model All-in-One
A single multi-modal diffusion transformer trained with masked learning performs pose estimation, depth prediction, and novel view synthesis in one model, reporting SOTA pose and NVS numbers.
-
Zero-1-to-G: Taming Pretrained 2D Diffusion Model for Direct 3D Generation
A single image can be turned into a 3D Gaussian splat model by fine-tuning a pretrained 2D diffusion model to output decomposed multi-view splatter attribute images.
-
Few-step Flow for 3D Generation via Marginal-Data Transport Distillation
MDT-dist distills a pretrained 3D flow model into a 1-2 step generator using velocity matching plus velocity distillation, cutting TRELLIS inference from 6.1s to 0.68s while approximately preserving generation quality.
-
DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization
Applying pairwise direct preference optimization to score distillation makes text-to-3D outputs better aligned with human preferences and more controllable.
-
InsTex: Indoor Scenes Stylized Texture Synthesis
InsTex generates style-consistent textures for indoor 3D scenes using a coarse-to-fine diffusion pipeline with global image guidance, reporting faster and higher-scoring results than four baselines.
-
Instructive3D: Editing Large Reconstruction Models with Text Instructions
A text-conditioned diffusion adapter operating on the triplane latents of a frozen large reconstruction model enables natural-language editing of generated 3D objects.
-
Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution
DDR-SR routes each real-world low-resolution image to one of two diffusion experts based on a high-frequency-loss difficulty score, using a low-compression VAE for hard images and a high-compression VAE for easy image...
-
EgoAnimate: Generating Human Animations from Egocentric top-down Views
EgoAnimate synthesizes a frontal T-pose image from an egocentric top-down photo using a fine-tuned Stable Diffusion model, then animates it with off-the-shelf image-to-motion methods to produce an animatable avatar.
-
DreamComposer++: Empowering Diffusion Models with Multi-View Conditions for 3D Content Generation
A multi-view conditioning framework that improves controllable novel view synthesis and 3D reconstruction by injecting fused 3D latents into frozen image and video diffusion models.
-
Acquiring and Adapting Priors for Novel Tasks via Neural Meta-Architectures
A meta-learning dissertation showing that distributed memory and hypernetworks can adapt to new tasks with few samples, applied to image classification, text-to-3D generation, and molecular binding prediction, with th...
Discussion (0). Continue with ORCID to comment.