REVIEW 16 cited by
LucidDreamer: Towards High-Fidelity Text-to-3D Generation via Interval Score Matching
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The recent advancements in text-to-3D generation mark a significant milestone in generative models, unlocking new possibilities for creating imaginative 3D assets across various real-world scenarios. While recent advancements in text-to-3D generation have shown promise, they often fall short in rendering detailed and high-quality 3D models. This problem is especially prevalent as many methods base themselves on Score Distillation Sampling (SDS). This paper identifies a notable deficiency in SDS, that it brings inconsistent and low-quality updating direction for the 3D model, causing the over-smoothing effect. To address this, we propose a novel approach called Interval Score Matching (ISM). ISM employs deterministic diffusing trajectories and utilizes interval-based score matching to counteract over-smoothing. Furthermore, we incorporate 3D Gaussian Splatting into our text-to-3D generation pipeline. Extensive experiments show that our model largely outperforms the state-of-the-art in quality and training efficiency.
Forward citations
Cited by 16 Pith papers
-
CORGI: Consistency-Aware 3D Dog Reconstruction from a Single Image in the Wild
CORGI reconstructs high-fidelity, animatable 3D dogs from a single in-the-wild image via canonical orbital generation, deformable 3DGS anchored to D-SMAL, and self-supervised generative repair, without 3D supervision.
-
ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space
A unified pipeline lifts any text/image/video input into a Spatial Generative Primitive, explores it with 3D-consistent panoramic video, and reconstructs photorealistic 3DGS worlds with stronger rich-input fidelity th...
-
MultiDreamer3D: Multi-concept 3D Customization with Concept-Aware Diffusion Guidance
MultiDreamer3D generates 3D scenes containing multiple personalized concepts by laying out bounding boxes, seeding coarse point clouds, and refining Gaussian splats with concept-aware diffusion guidance.
-
Consistent Flow Distillation for Text-to-3D Generation
Consistent Flow Distillation (CFD) guides 3D generation by denoising rendered views with a noise field that is consistent across camera views on the object surface.
-
Dora: Sampling and Benchmarking for 3D Shape Variational Auto-Encoders
Dora-VAE uses sharp-edge sampling plus dual cross-attention to match XCube-level reconstruction with 1,280 latent codes; Dora-bench adds complexity tiers and a sharp normal error metric.
-
GaussianPainter: Painting Point Cloud into 3D Gaussians with Normal Guidance
GaussianPainter produces 3D Gaussians from a point cloud and reference image in one forward pass by constraining Gaussian rotations with predicted surface normals.
-
PRM: Photometric Stereo based Large Reconstruction Model
PRM uses photometric-stereo-style rendered images as both input and supervision, with mesh-based differentiable PBR, to reconstruct 3D meshes with finer local details and more robustness to complex appearances.
-
Text-to-3D Generation by 2D Editing
GE3D generates 3D objects by aligning latents across multi-step 2D diffusion editing trajectories, replacing the single-step SDS noise comparison and yielding more photorealistic results.
-
AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers
AC3D improves camera control in video diffusion transformers by conditioning only early denoising steps and the first 8 of 32 blocks, and by adding 20K static-camera dynamic videos to training.
-
DetailGen3D: Generative 3D Geometry Enhancement via Data-Dependent Flow
DetailGen3D refines coarse 3D geometry into detailed geometry by learning a direct latent-space flow from coarse to fine shapes, guided by an input image.
-
Rethinking Score Distilling Sampling for 3D Editing and Generation
UDS unifies text-to-3D generation and 3D editing with a single score-distillation gradient formula that replaces noise with clean-latent estimates, reporting higher CLIP scores and user preference than prior SDS variants.
-
DiMeR: Disentangled Mesh Reconstruction Model
DiMeR reconstructs 3D meshes from sparse views by feeding normal maps into the geometry branch and RGB into a separate texture branch, cutting Chamfer Distance by up to 31.7% on GSO when ground-truth normals are used.
-
Towards Generalized and Training-Free Text-Guided Semantic Manipulation
GTF is a training-free, projection-based noise composition rule that enables text-driven addition, removal, and style transfer in diffusion models across image, video, and 3D generation.
-
Bench2Drive-R: Turning Real World Data into Reactive Closed-Loop Autonomous Driving Benchmark by Generative Model
A reactive closed-loop driving simulator that uses a diffusion renderer with retrieval from real recordings, plus a nuPlan behavioral controller, to generate sensor images in response to an end-to-end driving model's actions.
-
DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation
A pipeline that generates editable 3D scenes from natural language by combining LLM-based layout planning, multi-timestep diffusion distillation, and staged camera sampling.
-
Dive3D: Diverse Distillation-based Text-to-3D Generation via Score Implicit Matching
Dive3D shows that replacing KL divergence with score implicit matching in text-to-3D distillation, together with a reward term, produces more diverse and higher-fidelity 3D assets than SDS and ProlificDreamer baselines.
Discussion (0). Continue with ORCID to comment.