REVIEW 9 cited by
MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware Diffusion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
This paper introduces MVDiffusion, a simple yet effective method for generating consistent multi-view images from text prompts given pixel-to-pixel correspondences (e.g., perspective crops from a panorama or multi-view images given depth maps and poses). Unlike prior methods that rely on iterative image warping and inpainting, MVDiffusion simultaneously generates all images with a global awareness, effectively addressing the prevalent error accumulation issue. At its core, MVDiffusion processes perspective images in parallel with a pre-trained text-to-image diffusion model, while integrating novel correspondence-aware attention layers to facilitate cross-view interactions. For panorama generation, while only trained with 10k panoramas, MVDiffusion is able to generate high-resolution photorealistic images for arbitrary texts or extrapolate one perspective image to a 360-degree view. For multi-view depth-to-image generation, MVDiffusion demonstrates state-of-the-art performance for texturing a scene mesh.
Forward citations
Cited by 9 Pith papers
-
DMS:Diffusion-Based Multi-Baseline Stereo Generation for Improving Self-Supervised Depth Estimation
DMS fine-tunes Stable Diffusion to synthesize left-left, right-right, and center views from stereo images, then uses these synthetic views in a per-pixel minimum warping loss to reduce outliers in self-supervised dept...
-
CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-Centric 3D Scene Generation
CGGS generates viewpoint-consistent, text-aligned ego-centric 3D scenes via consistency-augmented multi-view diffusion, flow-guided layout initialization, and mutual-information depth-refined Gaussian optimization.
-
PixGS: Pixel-Space Diffusion for Direct 3D Gaussian Splat Generation
PixGS is a single-stage pixel-space diffusion model that directly produces high-quality 3D Gaussian Splats from text or images in ~1s, outperforming multi-stage latent methods on standard benchmarks.
-
Matrix3D: Large Photogrammetry Model All-in-One
A single multi-modal diffusion transformer trained with masked learning performs pose estimation, depth prediction, and novel view synthesis in one model, reporting SOTA pose and NVS numbers.
-
Multi-view Image Diffusion via Coordinate Noise and Fourier Attention
A diffusion method uses coordinate noise, time-dependent Fourier attention, and a cross-attention loss to improve multi-view consistency in generated images.
-
TiP4GEN: Text to Immersive Panorama 4D Scene Generation
TiP4GEN generates motion-rich, geometry-consistent 360-degree 4D scenes from a global text prompt plus four local perspective prompts, using a dual-branch video diffusion model with bidirectional cross-attention and a...
-
InsTex: Indoor Scenes Stylized Texture Synthesis
InsTex generates style-consistent textures for indoor 3D scenes using a coarse-to-fine diffusion pipeline with global image guidance, reporting faster and higher-scoring results than four baselines.
-
Pragmatist: Multiview Conditional Diffusion Models for High-Fidelity 3D Reconstruction from Unposed Sparse Views
Pragmatist turns sparse unposed photos of an object into a high-fidelity 3D mesh by generating consistent canonical views with a diffusion model, reconstructing a triplane mesh, then refining camera poses and texture ...
-
Physical Informed Driving World Model
DrivePhysica adds coordinate alignment, 3D instance flow, and box-coordinate guidance to a diffusion world model, achieving state-of-the-art FID/FVD on nuScenes and improving StreamPETR NDS by 3.6 points when mixed wi...
Discussion (0). Continue with ORCID to comment.