Pith. sign in

REVIEW 2 cited by

HexaGen3D: StableDiffusion is just one step away from Fast and Diverse Text-to-3D Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.07727 v1 pith:YFYHBAOA submitted 2024-01-15 cs.CV

classification cs.CV
keywords hexagen3dapproachassetsdiversegenerationhigh-qualityobjectspretrained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the latest remarkable advances in generative modeling, efficient generation of high-quality 3D assets from textual prompts remains a difficult task. A key challenge lies in data scarcity: the most extensive 3D datasets encompass merely millions of assets, while their 2D counterparts contain billions of text-image pairs. To address this, we propose a novel approach which harnesses the power of large, pretrained 2D diffusion models. More specifically, our approach, HexaGen3D, fine-tunes a pretrained text-to-image model to jointly predict 6 orthographic projections and the corresponding latent triplane. We then decode these latents to generate a textured mesh. HexaGen3D does not require per-sample optimization, and can infer high-quality and diverse objects from textual prompts in 7 seconds, offering significantly better quality-to-latency trade-offs when comparing to existing approaches. Furthermore, HexaGen3D demonstrates strong generalization to new objects or compositions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  2. Edit360: 2D Image Edits to 3D Assets from Any Angle

    cs.GR 2025-06 conditional novelty 6.0 of 10

    A training-free method that propagates a 2D edit applied at any chosen viewpoint across a full 360-degree orbit by fusing anchor-view and front-view video-diffusion trajectories.

Pith tools