REVIEW 12 cited by
Cycle3D: High-quality and Consistent Image-to-3D Generation via Generation-Reconstruction Cycle
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent 3D large reconstruction models typically employ a two-stage process, including first generate multi-view images by a multi-view diffusion model, and then utilize a feed-forward model to reconstruct images to 3D content.However, multi-view diffusion models often produce low-quality and inconsistent images, adversely affecting the quality of the final 3D reconstruction. To address this issue, we propose a unified 3D generation framework called Cycle3D, which cyclically utilizes a 2D diffusion-based generation module and a feed-forward 3D reconstruction module during the multi-step diffusion process. Concretely, 2D diffusion model is applied for generating high-quality texture, and the reconstruction model guarantees multi-view consistency.Moreover, 2D diffusion model can further control the generated content and inject reference-view information for unseen views, thereby enhancing the diversity and texture consistency of 3D generation during the denoising process. Extensive experiments demonstrate the superior ability of our method to create 3D content with high-quality and consistency compared with state-of-the-art baselines.
Forward citations
Cited by 12 Pith papers
-
Hypothesis-and-Refinement Learning of Organic Structures from Multimodal Spectroscopic Data
A two-stage AI pipeline — spectral hypothesis generation followed by mass-constrained molecular refinement — reconstructs organic structures from multimodal spectra, with 93.8% top-1 accuracy on simulated QM9 data and...
-
E-4DGS: High-Fidelity Dynamic Reconstruction from the Multi-view Event Cameras
E-4DGS is a deformable 3D Gaussian Splatting method that reconstructs dynamic scenes directly from multi-view event camera streams, outperforming event-to-image baseline approaches.
-
OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
A benchmark and five-million-clip dataset for evaluating and training subject-to-video generation models, with three new metrics for subject consistency, naturalness, and text alignment.
-
GS2E: Gaussian Splatting is an Effective Data Generator for Event Stream Generation
A pipeline that turns sparse multi-view RGB images into a claimed 1,150-scene synthetic event dataset using 3D Gaussian Splatting rendering plus a stochastic event simulator.
-
EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion
EchoVideo preserves identity in generated human videos by pre-fusing face, image, and text features, then training with stochastic shallow-feature dropout to reduce copy-paste artifacts.
-
Ingredients: Blending Custom Photos with Video Diffusion Transformers
Ingredients adds a mask-supervised identity router to a video diffusion transformer, enabling multi-person videos from a few reference photos without per-identity fine-tuning.
-
Navigating Chemical-Linguistic Sharing Space with Heterogeneous Molecular Encoding
A multi-view molecular encoder with fragment-based chain-of-thought improves text-to-molecule and molecule-to-text generation, supported by a new one-million-molecule conditional design dataset.
-
DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses
A two-stage diffusion pipeline that enriches 2D pose guidance with generated depth and normal maps to achieve state-of-the-art human image animation.
-
Identity-Preserving Text-to-Video Generation by Frequency Decomposition
ConsisID generates identity-preserving videos by injecting low-frequency facial features into shallow layers and high-frequency identity features into attention blocks of a DiT video model.
-
Fancy123: One Image to High-Quality 3D Mesh Generation via Plug-and-Play Deformation
Fancy123 refines single-image-to-3D meshes by deforming multiview images and then the mesh itself, then unprojecting clear image colors onto the surface.
-
AE-NeRF: Augmenting Event-Based Neural Radiance Fields for Non-ideal Conditions and Larger Scene
AE-NeRF jointly optimizes camera poses and an event-based NeRF with a proposal network and four event-specific losses, improving novel view synthesis under noisy poses and non-uniform motion.
-
Hierarchical Banzhaf Interaction for General Video-Language Representation Learning
HBI V2 models video-text alignment as a cooperative game with Hierarchical Banzhaf Interaction plus single/cross-modal representation fusion, improving retrieval, QA, and captioning benchmarks.
Discussion (0). Continue with ORCID to comment.