Pith. sign in

REVIEW 4 cited by

Phidias: A Generative Model for Creating 3D Content from Text, Image, and 3D Conditions with Reference-Augmented Diffusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.11406 v1 pith:5DPQBL34 submitted 2024-09-17 cs.CV

classification cs.CV
keywords modelgenerationimagereferencephidiasconditionsdiffusionexisting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In 3D modeling, designers often use an existing 3D model as a reference to create new ones. This practice has inspired the development of Phidias, a novel generative model that uses diffusion for reference-augmented 3D generation. Given an image, our method leverages a retrieved or user-provided 3D reference model to guide the generation process, thereby enhancing the generation quality, generalization ability, and controllability. Our model integrates three key components: 1) meta-ControlNet that dynamically modulates the conditioning strength, 2) dynamic reference routing that mitigates misalignment between the input image and 3D reference, and 3) self-reference augmentations that enable self-supervised training with a progressive curriculum. Collectively, these designs result in a clear improvement over existing methods. Phidias establishes a unified framework for 3D generation using text, image, and 3D conditions with versatile applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level Control

    cs.GR 2026-07 conditional novelty 6.5 of 10

    Automatic view scheduling via a directed generation graph plus object-level identity and adherence conditioning enables high-quality outdoor 3DGS scenes from arbitrary input geometry without user camera paths.

  2. Neural LightRig: Unlocking Accurate Object Normal and Material Estimation with Multi-Light Diffusion

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A diffusion model creates multi-light images from one photo, and a U-Net uses them to predict object normals and PBR materials, improving single-image inverse rendering on synthetic benchmarks.

  3. HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A staged pipeline generates layered, mesh-based 3D worlds from text or images by combining panoramic diffusion, semantic layer decomposition, and video-based expansion.

  4. Material Anything: Generating Materials for Any 3D Object via Diffusion

    cs.CV 2024-11 conditional novelty 5.0 of 10

    Material Anything is a unified diffusion pipeline that generates PBR material maps (albedo, roughness, metallic, bump) for arbitrary 3D meshes using confidence masks to handle varying texture and lighting conditions.

Pith tools