Pith. sign in

REVIEW 16 cited by

LucidDreamer: Towards High-Fidelity Text-to-3D Generation via Interval Score Matching

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.11284 v3 pith:2XB2RDJ5 submitted 2023-11-19 cs.CV cs.GRcs.MM

classification cs.CVcs.GRcs.MM
keywords generationscoretext-to-3dmatchingadvancementsintervalmodelmodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The recent advancements in text-to-3D generation mark a significant milestone in generative models, unlocking new possibilities for creating imaginative 3D assets across various real-world scenarios. While recent advancements in text-to-3D generation have shown promise, they often fall short in rendering detailed and high-quality 3D models. This problem is especially prevalent as many methods base themselves on Score Distillation Sampling (SDS). This paper identifies a notable deficiency in SDS, that it brings inconsistent and low-quality updating direction for the 3D model, causing the over-smoothing effect. To address this, we propose a novel approach called Interval Score Matching (ISM). ISM employs deterministic diffusing trajectories and utilizes interval-based score matching to counteract over-smoothing. Furthermore, we incorporate 3D Gaussian Splatting into our text-to-3D generation pipeline. Extensive experiments show that our model largely outperforms the state-of-the-art in quality and training efficiency.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CORGI: Consistency-Aware 3D Dog Reconstruction from a Single Image in the Wild

    cs.CV 2026-07 unverdicted novelty 7.0 of 10

    CORGI reconstructs high-fidelity, animatable 3D dogs from a single in-the-wild image via canonical orbital generation, deformable 3DGS anchored to D-SMAL, and self-supervised generative repair, without 3D supervision.

  2. ABot-3DWorld 0: A Universal World Model to Explore Any 3D Space

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    A unified pipeline lifts any text/image/video input into a Spatial Generative Primitive, explores it with 3D-consistent panoramic video, and reconstructs photorealistic 3DGS worlds with stronger rich-input fidelity th...

  3. MultiDreamer3D: Multi-concept 3D Customization with Concept-Aware Diffusion Guidance

    cs.CV 2025-01 conditional novelty 6.0 of 10

    MultiDreamer3D generates 3D scenes containing multiple personalized concepts by laying out bounding boxes, seeding coarse point clouds, and refining Gaussian splats with concept-aware diffusion guidance.

  4. Consistent Flow Distillation for Text-to-3D Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Consistent Flow Distillation (CFD) guides 3D generation by denoising rendered views with a noise field that is consistent across camera views on the object surface.

  5. Dora: Sampling and Benchmarking for 3D Shape Variational Auto-Encoders

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Dora-VAE uses sharp-edge sampling plus dual cross-attention to match XCube-level reconstruction with 1,280 latent codes; Dora-bench adds complexity tiers and a sharp normal error metric.

  6. GaussianPainter: Painting Point Cloud into 3D Gaussians with Normal Guidance

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GaussianPainter produces 3D Gaussians from a point cloud and reference image in one forward pass by constraining Gaussian rotations with predicted surface normals.

  7. PRM: Photometric Stereo based Large Reconstruction Model

    cs.CV 2024-12 conditional novelty 6.0 of 10

    PRM uses photometric-stereo-style rendered images as both input and supervision, with mesh-based differentiable PBR, to reconstruct 3D meshes with finer local details and more robustness to complex appearances.

  8. Text-to-3D Generation by 2D Editing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    GE3D generates 3D objects by aligning latents across multi-step 2D diffusion editing trajectories, replacing the single-step SDS noise comparison and yielding more photorealistic results.

  9. AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers

    cs.CV 2024-11 conditional novelty 6.0 of 10

    AC3D improves camera control in video diffusion transformers by conditioning only early denoising steps and the first 8 of 32 blocks, and by adding 20K static-camera dynamic videos to training.

  10. DetailGen3D: Generative 3D Geometry Enhancement via Data-Dependent Flow

    cs.CV 2024-11 conditional novelty 6.0 of 10

    DetailGen3D refines coarse 3D geometry into detailed geometry by learning a direct latent-space flow from coarse to fine shapes, guided by an input image.

  11. Rethinking Score Distilling Sampling for 3D Editing and Generation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    UDS unifies text-to-3D generation and 3D editing with a single score-distillation gradient formula that replaces noise with clean-latent estimates, reporting higher CLIP scores and user preference than prior SDS variants.

  12. DiMeR: Disentangled Mesh Reconstruction Model

    cs.CV 2025-04 conditional novelty 5.0 of 10

    DiMeR reconstructs 3D meshes from sparse views by feeding normal maps into the geometry branch and RGB into a separate texture branch, cutting Chamfer Distance by up to 31.7% on GSO when ground-truth normals are used.

  13. Towards Generalized and Training-Free Text-Guided Semantic Manipulation

    cs.CV 2025-04 conditional novelty 5.0 of 10

    GTF is a training-free, projection-based noise composition rule that enables text-driven addition, removal, and style transfer in diffusion models across image, video, and 3D generation.

  14. Bench2Drive-R: Turning Real World Data into Reactive Closed-Loop Autonomous Driving Benchmark by Generative Model

    cs.RO 2024-12 conditional novelty 5.0 of 10

    A reactive closed-loop driving simulator that uses a diffusion renderer with retrieval from real recordings, plus a nuPlan behavioral controller, to generate sensor images in response to an end-to-end driving model's actions.

  15. DreamScene: 3D Gaussian-based End-to-end Text-to-3D Scene Generation

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A pipeline that generates editable 3D scenes from natural language by combining LLM-based layout planning, multi-timestep diffusion distillation, and staged camera sampling.

  16. Dive3D: Diverse Distillation-based Text-to-3D Generation via Score Implicit Matching

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Dive3D shows that replacing KL divergence with score implicit matching in text-to-3D distillation, together with a reward term, produces more diverse and higher-fidelity 3D assets than SDS and ProlificDreamer baselines.

Pith tools