Pith. sign in

REVIEW 8 cited by

SVDiff: Compact Parameter Space for Diffusion Fine-Tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.11305 v4 pith:G2R2IOCW submitted 2023-03-20 cs.CV

classification cs.CV
keywords diffusionexistingmodelscompactcomparedfine-tuninggenerationimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models have achieved remarkable success in text-to-image generation, enabling the creation of high-quality images from text prompts or other modalities. However, existing methods for customizing these models are limited by handling multiple personalized subjects and the risk of overfitting. Moreover, their large number of parameters is inefficient for model storage. In this paper, we propose a novel approach to address these limitations in existing text-to-image diffusion models for personalization. Our method involves fine-tuning the singular values of the weight matrices, leading to a compact and efficient parameter space that reduces the risk of overfitting and language drifting. We also propose a Cut-Mix-Unmix data-augmentation technique to enhance the quality of multi-subject image generation and a simple text-based image editing framework. Our proposed SVDiff method has a significantly smaller model size compared to existing methods (approximately 2,200 times fewer parameters compared with vanilla DreamBooth), making it more practical for real-world applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. R^2MoE: Redundancy-Removal Mixture of Experts for Lifelong Concept Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    R2MoE adds per-concept LoRA experts with routing distillation and expert pruning, reporting 0.19% forgetting and 15.2M added parameters on CustomConcept101.

  2. UnZipLoRA: Separating Content and Style from a Single Image

    cs.CV 2024-12 conditional novelty 6.0 of 10

    From a single image, UnZipLoRA jointly learns a content LoRA and a style LoRA that can be used separately or combined by direct addition.

  3. PersonaCraft: Personalized and Controllable Full-Body Multi-Human Scene Generation Using Occlusion-Aware 3D-Conditioned Diffusion

    cs.CV 2024-11 conditional novelty 6.0 of 10

    PersonaCraft adds SMPLx depth and normal conditioning, occlusion boundary enhancement, and occlusion-aware classifier-free guidance to diffusion models, enabling controllable multi-person images that preserve both fac...

  4. Sharp Generalization Bounds for Foundation Models with Asymmetric Randomized Low-Rank Adapters

    stat.ML 2025-06 reject novelty 5.0 of 10

    The paper proves an upper bound of about sqrt(r/N) on the LoRA generalization gap and claims a matching lower bound, but both proofs contain structural gaps.

  5. Direct Preference Optimization-Enhanced Multi-Guided Diffusion Model for Traffic Scenario Generation

    cs.LG 2025-02 reject novelty 5.0 of 10

    MuDi-Pro fine-tunes a multi-guided diffusion transformer with DPO using guidance-score preferences to improve controllability of traffic scenario generation on nuScenes.

  6. AnyStory: Towards Unified Single and Multiple Subject Personalization in Text-to-Image Generation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    AnyStory introduces a unified feed-forward approach for single and multi-subject text-to-image personalization using a simplified ReferenceNet and CLIP encoder, plus a decoupled instance-aware router.

  7. ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A feed-forward multi-concept video customization model that fuses each concept image with its text label and injects the composite embeddings through a separate cross-attention layer, avoiding test-time optimization.

  8. Stage-Aware Adaptation and Distribution Calibration for Subject-Driven Personalized Text-to-Image Generation

    cs.CV 2026-07 conditional novelty 3.0 of 10

    Stage-aware low-rank scaling and distribution-calibrated candidate selection improve identity consistency in personalized image generation but reduce output diversity.

Pith tools