Pith. sign in

REVIEW 13 cited by

Generative AI meets 3D: A Survey on Text-to-3D in AIGC Era

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.06131 v4 pith:ORZSIP7W submitted 2023-05-10 cs.CV

classification cs.CV
keywords text-to-3dgenerationfieldincludingresearchaigccontentdata
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Generative AI has made significant progress in recent years, with text-guided content generation being the most practical as it facilitates interaction between human instructions and AI-generated content (AIGC). Thanks to advancements in text-to-image and 3D modeling technologies, like neural radiance field (NeRF), text-to-3D has emerged as a nascent yet highly active research field. Our work conducts a comprehensive survey on this topic and follows up on subsequent research progress in the overall field, aiming to help readers interested in this direction quickly catch up with its rapid development. First, we introduce 3D data representations, including both Structured and non-Structured data. Building on this pre-requisite, we introduce various core technologies to achieve satisfactory text-to-3D results. Additionally, we present mainstream baselines and research directions in recent text-to-3D technology, including fidelity, efficiency, consistency, controllability, diversity, and applicability. Furthermore, we summarize the usage of text-to-3D technology in various applications, including avatar generation, texture generation, scene generation and 3D editing. Finally, we discuss the agenda for the future development of text-to-3D.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  2. A11yShape: AI-Assisted 3-D Modeling for Blind and Low-Vision Programmers

    cs.HC 2025-08 conditional novelty 6.0 of 10

    A11yShape enables blind and low-vision programmers to create and modify 3-D models through an AI-assisted, code-based system, as demonstrated in a four-participant study.

  3. "You'll Be Alice Adventuring in Wonderland!" Processes, Challenges, and Opportunities of Creating Animated Virtual Reality Stories

    cs.HC 2025-02 conditional novelty 6.0 of 10

    A 21-creator interview study identifies ten common stages and nine unique challenges in creating animated VR stories, with story-driven and visual-driven workflow types.

  4. CreepyCoCreator? Investigating AI Representation Modes for 3D Object Co-Creation in Virtual Reality

    cs.HC 2025-02 conditional novelty 6.0 of 10

    A Wizard-of-Oz VR study shows that embodiment, highlighting, and incremental visualization each shape how users perceive an AI co-creator, with embodiment increasing perceived partnership and contribution.

  5. MagicNaming: Consistent Identity Generation by Finding a "Name Space" in T2I Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    An image encoder maps any face to a 'name embedding' that, when prepended to a text prompt, makes an SDXL model generate consistent identities for arbitrary people without fine-tuning.

  6. MARVEL-40M+: Multi-Level Visual Elaboration for High-Fidelity Text-to-3D Content Creation

    cs.CV 2024-11 conditional novelty 6.0 of 10

    MARVEL-40M+ provides multi-level captions for over 8.9 million 3D assets and a two-stage text-to-3D pipeline that generates textured meshes in 15 seconds.

  7. PlantDreamer: Achieving Realistic 3D Plant Models with Diffusion-Guided Gaussian Splatting

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A diffusion-guided Gaussian splatting pipeline generates realistic 3D plants from L-system meshes or point clouds and beats GaussianDreamer on masked PSNR for bean, kale and mint.

  8. From Air to Wear: Personalized 3D Digital Fashion with AR/VR Immersive 3D Sketching

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A VR-sketch-conditioned diffusion model, trained in three stages with curriculum learning and a new 969-pair dataset, generates plausible 3D garments from freehand 3D sketches.

  9. DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Applying pairwise direct preference optimization to score distillation makes text-to-3D outputs better aligned with human preferences and more controllable.

  10. Toward Scene Graph and Layout Guided Complex 3D Scene Generation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    GraLa3D builds a graph with single-object nodes and super-nodes plus layout boxes, then generates 3D Gaussian scenes from those structures to preserve spatial and interaction relations.

  11. Reconstructing 4D Spatial Intelligence: A Survey

    cs.CV 2025-07 accept novelty 4.0 of 10

    A review that classifies 4D scene reconstruction methods into five progressive levels: low-level cues, scene components, dynamic scenes, interactions, and physics.

  12. Making Physical Objects with Generative AI and Robotic Assembly: Considering Fabrication Constraints, Sustainability, Time, Functionality, and Accessibility

    cs.RO 2025-04 conditional novelty 4.0 of 10

    A text-to-3D model plus a robot that stacks magnetic blocks can make simple objects in minutes, but the claims rest on a small, anecdotal case study.

  13. 3D Scene Generation: A Survey

    cs.CV 2025-05 conditional

    The paper surveys 3D scene generation and organizes methods into four paradigms, with datasets, evaluation metrics, applications, and future directions.

Pith tools