Pith. sign in

REVIEW 20 cited by

MV-Adapter: Multi-view Consistent Image Generation Made Easy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.03632 v1 pith:MA3736K3 submitted 2024-12-04 cs.CV

classification cs.CV
keywords generationimagemodelsmulti-viewmv-adapteradapterknowledgepre-trained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing multi-view image generation methods often make invasive modifications to pre-trained text-to-image (T2I) models and require full fine-tuning, leading to (1) high computational costs, especially with large base models and high-resolution images, and (2) degradation in image quality due to optimization difficulties and scarce high-quality 3D data. In this paper, we propose the first adapter-based solution for multi-view image generation, and introduce MV-Adapter, a versatile plug-and-play adapter that enhances T2I models and their derivatives without altering the original network structure or feature space. By updating fewer parameters, MV-Adapter enables efficient training and preserves the prior knowledge embedded in pre-trained models, mitigating overfitting risks. To efficiently model the 3D geometric knowledge within the adapter, we introduce innovative designs that include duplicated self-attention layers and parallel attention architecture, enabling the adapter to inherit the powerful priors of the pre-trained models to model the novel 3D knowledge. Moreover, we present a unified condition encoder that seamlessly integrates camera parameters and geometric information, facilitating applications such as text- and image-based 3D generation and texturing. MV-Adapter achieves multi-view generation at 768 resolution on Stable Diffusion XL (SDXL), and demonstrates adaptability and versatility. It can also be extended to arbitrary view generation, enabling broader applications. We demonstrate that MV-Adapter sets a new quality standard for multi-view image generation, and opens up new possibilities due to its efficiency, adaptability and versatility.

Discussion (0). Sign in to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AdaptSplat: Adapting Vision Foundation Models for Feed-Forward 3D Gaussian Splatting

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    AdaptSplat adds a Frequency-Preserving Adapter to vision foundation models to boost high-frequency fidelity and cross-domain performance in feed-forward 3D Gaussian Splatting.

  2. Towards Realistic and Consistent Orbital Video Generation via 3D Foundation Priors

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    A video generation approach conditions a base model with multi-scale 3D latent features and a cross-attention adapter to produce geometrically realistic and consistent orbital videos from one image.

  3. VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space

    cs.CV 2025-08 conditional novelty 7.0 of 10

    A training-free 3D editing method that inverts a source asset into TRELLIS latent space and replaces latents plus attention K/V tokens in unedited regions during re-denosing.

  4. MV-RAG: Retrieval Augmented Multiview Diffusion

    cs.CV 2025-08 conditional novelty 7.0 of 10

    A retrieval-augmented multiview diffusion model conditions on web images to generate 3D-consistent views of rare concepts, trained with a hybrid 3D/2D objective and evaluated on a new OOD benchmark.

  5. PacTure: Efficient PBR Texture Generation on Packed Views with Visual Autoregressive Models

    cs.CV 2025-05 unverdicted novelty 7.0 of 10

    PacTure uses view packing and next-scale autoregressive prediction to generate consistent multi-view PBR textures faster than prior sequential or cross-attention methods.

  6. EditFlow3D: Automated Local Editing of 3D Assets with Trajectory Preservation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Mask-guided differential flow with a soft preservation loss enables training-free local 3D editing that keeps unedited regions close to the source asset.

  7. PixGS: Pixel-Space Diffusion for Direct 3D Gaussian Splat Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    PixGS is a single-stage pixel-space diffusion model that directly produces high-quality 3D Gaussian Splats from text or images in ~1s, outperforming multi-stage latent methods on standard benchmarks.

  8. PixGS: Pixel-Space Diffusion for Direct 3D Gaussian Splat Generation

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    A single-stage pixel-space diffusion model for direct 3D Gaussian Splat generation that bypasses latent compression and adds geometric supervisions to outperform prior multi-stage methods.

  9. ReplicateAnyScene: Zero-Shot Video-to-3D Composition via Textual-Visual-Spatial Alignment

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    ReplicateAnyScene performs fully automated zero-shot video-to-compositional-3D reconstruction by cascading alignments of generic priors from vision foundation models across textual, visual, and spatial dimensions.

  10. TInR: Exploring Tool-Internalized Reasoning in Large Language Models

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    TInR-U internalizes tool knowledge into LLMs via bidirectional alignment, supervised fine-tuning, and reinforcement learning, outperforming standard tool-integrated reasoning in both in-domain and out-of-domain evaluations.

  11. SegviGen: Repurposing 3D Generative Model for Part Segmentation

    cs.CV 2026-03 unverdicted novelty 6.0 of 10

    SegviGen shows pretrained 3D generative models can be repurposed for part segmentation via voxel colorization, beating prior methods by 40% interactively and 15% on full segmentation using only 0.32% of labeled data.

  12. ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models

    cs.CV 2025-12 conditional novelty 6.0 of 10

    A masked discrete-diffusion transformer generates multiple consistent object views from a single image or text, reporting the best average PSNR/SSIM/LPIPS on GSO and 3D-FUTURE.

  13. Restore3D: Breathing Life into Broken Objects with Shape and Texture Restoration

    cs.CV 2026-07 unverdicted novelty 5.0 of 10

    Restore3D restores shape and texture of broken 3D objects via multi-view image refinement with a Mask Self-Perceiver and coarse-to-fine mesh reconstruction, outperforming baselines on synthetic and real benchmarks.

  14. Flex4DHuman: Flexible Multi-view Video Diffusion for 4D Human Reconstruction

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    A multi-view video diffusion model conditioned on relative camera poses via extended RoPE generates dense synchronized views from sparse inputs for 4D Gaussian splatting reconstruction, claiming SOTA results on human ...

  15. AdaptSplat: Adapting Vision Foundation Models for Feed-Forward 3D Gaussian Splatting

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    AdaptSplat adds a lightweight Frequency-Preserving Adapter to vision foundation models that extracts direction-aware high-frequency priors and integrates them via positional encodings and residual modulation to improv...

  16. DreamLifting: A Plug-in Module Lifting MV Diffusion Models for 3D Asset Generation

    cs.CV 2025-09 unverdicted novelty 5.0 of 10

    LGAA is a modular adapter framework that lifts multi-view diffusion models to produce 2D Gaussian Splats with PBR channels for high-quality relightable 3D mesh extraction using data-efficient finetuning on 69k instances.

  17. ObjFiller3D: Scaling 3D Object Inpainting to Dense Multi-View Consistency

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    ObjFiller3D jointly optimizes a dense 360-degree ring of views to inpaint 3D objects with cross-view-consistent textures, reporting higher PSNR and LPIPS than per-view baselines at much lower runtime.

  18. Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis

    cs.CV 2025-08 conditional novelty 5.0 of 10

    An MLLM-driven pipeline that composes 3D scenes from assets, optimizes them with multi-view VLM feedback, and renders videos, yielding synthetic data that modestly improves several 2D, 3D, and 4D generative baselines.

  19. TInR: Exploring Tool-Internalized Reasoning in Large Language Models

    cs.CL 2026-04 unverdicted novelty 4.0 of 10

    TInR-U internalizes tool knowledge into an LLM via bidirectional alignment, SFT warm-up, and RL, claiming better in- and out-of-domain tool reasoning without external docs.

  20. Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation

    cs.CV 2025-01 unverdicted novelty 4.0 of 10

    Hunyuan3D 2.0 scales flow-based diffusion transformers and texture synthesis models to generate high-resolution textured 3D assets that outperform prior state-of-the-art in geometry, alignment, and texture quality.

Pith tools