Pith. sign in

REVIEW 7 cited by

Score Jacobian Chaining: Lifting Pretrained 2D Diffusion Models for 3D Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.00774 v1 pith:G5M6RFBF submitted 2022-12-01 cs.CV cs.LG

classification cs.CVcs.LG
keywords diffusionmodelscorefieldgenerationgradientsjacobianmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A diffusion model learns to predict a vector field of gradients. We propose to apply chain rule on the learned gradients, and back-propagate the score of a diffusion model through the Jacobian of a differentiable renderer, which we instantiate to be a voxel radiance field. This setup aggregates 2D scores at multiple camera viewpoints into a 3D score, and repurposes a pretrained 2D model for 3D data generation. We identify a technical challenge of distribution mismatch that arises in this application, and propose a novel estimation mechanism to resolve it. We run our algorithm on several off-the-shelf diffusion image generative models, including the recently released Stable Diffusion trained on the large-scale LAION dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos

    cs.CV 2024-12 conditional novelty 7.0 of 10

    A diffusion model trained on 1 million 360-degree videos synthesizes novel views with camera translation and enables 3D reconstruction from a single image.

  2. ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    ELSA3D introduces elastic semantic anchoring via sparse anchor tokens and a scale-aware octree tokenizer to unify 3D generation and captioning at reduced computational cost.

  3. Sharp-It: A Multi-view to Multi-view Diffusion Model for 3D Synthesis and Manipulation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Sharp-It fine-tunes a multi-view diffusion model to enhance low-quality Shap-E renderings into high-quality multi-view sets that can be reconstructed into detailed 3D assets.

  4. ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ShapeLLM-Omni unifies text, image, and 3D generation and understanding in one autoregressive LLM using discrete 3D tokens and a new 3D-Alpaca training dataset.

  5. SHaDe: Compact and Consistent Dynamic 3D Reconstruction via Tri-Plane Deformation and Latent Diffusion

    cs.CV 2025-05 reject novelty 5.0 of 10

    SHaDe combines explicit tri-plane deformation, SH attention rendering, and latent diffusion refinement to reconstruct dynamic 3D scenes from sparse multi-view images and reports state-of-the-art results on D-NeRF.

  6. Prometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene Generation

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A feed-forward system that generates object-level and scene-level 3D Gaussian scenes from text in about eight seconds by diffusing multi-view RGB-D latent codes and decoding them into pixel-aligned 3D Gaussians.

  7. CRAFT: Designing Creative and Functional 3D Objects

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A mesh deformation system that jointly optimizes semantic alignment with text or image prompts and body fit, producing body-aware 3D objects.

Pith tools