HIG enforces exact histogram constraints on diffusion-generated images by modeling the control task as an optimal transport problem and applying guidance transformations during sampling.
Instantstyle-plus: Style transfer with content-preserving in text-to-image generation
8 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Scheduled decreasing style injection across decoder layers and denoising timesteps, combined with ControlNet scheduling, expands the style-content tradeoff frontier and achieves 6.1% better ArtFID than StyleID across 28k images.
Raw CSD cosine is not a calibrated absolute style score for many artists; a discrimination-gap diagnostic flags the failures and CSLS on the frozen backbone corrects most of them.
DiLAST optimizes 3D latents via guidance from a 2D diffusion model to enable generalizable style transfer for OOD styles in 3D asset generation.
Rule-based and learning-based algorithms simplify dance motions to help novices learn more effectively while maintaining naturalness and style.
A training-free method modifies diffusion model sampling with differentiable Sliced 1-Wasserstein distance for color-conditional image generation.
Disco-LoRA proposes disentangling content-style and content-motion via dual-LoRA with statistical regularization to enable multi-concept video customization.
CraftGraffiti applies LoRA-tuned diffusion transformers followed by identity-augmented self-attention and CLIP-guided pose extension to generate graffiti while preserving facial features.
citing papers explorer
-
Histogram-constrained Image Generation
HIG enforces exact histogram constraints on diffusion-generated images by modeling the control task as an optimal transport problem and applying guidance transformations during sampling.
-
Scheduled Style Injection: Expanding the Style-Content Pareto Frontier in Training-Free Diffusion-based Style Transfer
Scheduled decreasing style injection across decoder layers and denoising timesteps, combined with ControlNet scheduling, expands the style-content tradeoff frontier and achieves 6.1% better ArtFID than StyleID across 28k images.
-
When Style Similarity Scores Fail: Diagnosing Raw CSD Cosine in Artist-Style Evaluation
Raw CSD cosine is not a calibrated absolute style score for many artists; a discrimination-gap diagnostic flags the failures and CSLS on the frozen backbone corrects most of them.
-
Structured 3D Latents Are Surprisingly Powerful: Unleashing Generalizable Style with 2D Diffusion
DiLAST optimizes 3D latents via guidance from a 2D diffusion model to enable generalizable style transfer for OOD styles in 3D asset generation.
-
Make it Simple, Make it Dance: Dance Motion Simplification to Support Novices' Dance Learning
Rule-based and learning-based algorithms simplify dance motions to help novices learn more effectively while maintaining naturalness and style.
-
Color Conditional Generation with Sliced Wasserstein Guidance
A training-free method modifies diffusion model sampling with differentiable Sliced 1-Wasserstein distance for color-conditional image generation.
-
Disco-LoRA: Disentangled Composition of Content, Style, and Motion for Multi-concept Video Customization
Disco-LoRA proposes disentangling content-style and content-motion via dual-LoRA with statistical regularization to enable multi-concept video customization.
-
CraftGraffiti: Exploring Human Identity with Custom Graffiti Art via Facial-Preserving Diffusion Models
CraftGraffiti applies LoRA-tuned diffusion transformers followed by identity-augmented self-attention and CLIP-guided pose extension to generate graffiti while preserving facial features.