REVIEW 55 cited by
ControlNeXt: Powerful and Efficient Control for Image and Video Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Diffusion models have demonstrated remarkable and robust abilities in both image and video generation. To achieve greater control over generated results, researchers introduce additional architectures, such as ControlNet, Adapters and ReferenceNet, to integrate conditioning controls. However, current controllable generation methods often require substantial additional computational resources, especially for video generation, and face challenges in training or exhibit weak control. In this paper, we propose ControlNeXt: a powerful and efficient method for controllable image and video generation. We first design a more straightforward and efficient architecture, replacing heavy additional branches with minimal additional cost compared to the base model. Such a concise structure also allows our method to seamlessly integrate with other LoRA weights, enabling style alteration without the need for additional training. As for training, we reduce up to 90% of learnable parameters compared to the alternatives. Furthermore, we propose another method called Cross Normalization (CN) as a replacement for Zero-Convolution' to achieve fast and stable training convergence. We have conducted various experiments with different base models across images and videos, demonstrating the robustness of our method.
Forward citations
Cited by 55 Pith papers
-
FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors
Interactive image editing can be cast as image-to-video generation: initializing from Stable Video Diffusion plus a new matching attention mechanism yields high-quality sketch, drag, and coarse-edit results with far l...
-
From Open Loop to Closed Loop: A Test-Time Iterative Optimization Framework for Reference-Consistent Image Generation
A training-free closed-loop PID controller iteratively corrects latent control signals so diffusion models stay consistent with ID, pose, or depth references better than matched open-loop sampling.
-
Appearance Pointers -- Multimodal Region Control of Diffusion Transformers
Appearance pointers are compact tokens that let a diffusion transformer apply text, image, or combined prompts to specific image regions in a single pass.
-
To Blend In, First Decouple: Rethinking Camouflage Image Generation via Context-Decoupled Representations
CamoDreamer generates camouflage images by decoupling foreground and background control in a diffusion model, reporting a 15.5-point FID gain over prior state of the art on LAKE-RED.
-
3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement
Presents a scene-adaptive 3D human image animation framework using ground-adaptive motion retargeting and viewpoint-adaptive latent fusion to control human and camera trajectories, claiming improvements on two benchmarks.
-
TerraDiT: Point-Conditioned Diffusion Transformer for Satellite Image Synthesis
GeoDiT is a point-conditioned diffusion transformer that generates satellite imagery from sparse labeled points and claims to beat existing remote sensing generators on FID and SSIM.
-
ANYPORTAL: Zero-Shot Consistent Video Background Replacement
A training-free video background replacement pipeline that keeps the foreground pixel-consistent by projecting refined latents through a deterministic reparameterization.
-
Extension of generalized KYP lemma: from LTI systems to LPV systems
The abstract claims a gKYP lemma extension for LPV systems via frequency-range enlargement, but the submitted full text is an unrelated video generation paper.
-
Multi-human Interactive Talking Dataset
The paper contributes a 12-hour multi-person conversational video dataset with pose and speaking annotations, plus a baseline model for generating full-body talking videos of two to four people.
-
DreamPainter: Image Background Inpainting for E-commerce Scenarios
DreamPainter introduces a two-stage diffusion framework trained on a new synthetic e-commerce dataset, DreamEcom-400K, that outperforms open-source inpainting baselines on background generation with text and reference...
-
TurboVSR: Fantastic Video Upscalers and Where to Find Them
TurboVSR uses a high-compression video autoencoder with factorized conditioning and non-uniform shortcut sampling to achieve near-state-of-the-art perceptual video super-resolution at roughly 100x lower compute cost.
-
OutDreamer: Video Outpainting with a Diffusion Transformer
OutDreamer couples a diffusion transformer with mask-driven self-attention and a latent alignment loss to outpaint videos in a zero-shot manner, exceeding prior zero-shot baselines on standard benchmarks.
-
PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation
PoseMaster produces a 3D character mesh from one image and a target 3D skeleton, preserving identity and pose in a single unified model, and it outperforms two-stage 2D-to-3D baselines on the VRoid pose canonicalizati...
-
Rethink Sparse Signals for Pose-guided Text-to-image Generation
SP-Ctrl improves pose-guided text-to-image generation with sparse poses by learning keypoint embeddings and supervising keypoint attention maps, nearly matching dense depth-based control.
-
DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers
A diffusion transformer model generates human-product demonstration videos from paired human and product images while preserving both identities through masked cross-attention and motion template guidance.
-
EF-VI: Enhancing End-Frame Injection for Video Inbetweening
EF-VI injects temporally expanded end-frame features into a transformer-based image-to-video diffusion model, improving video inbetweening quality over direct fine-tuning and bidirectional sampling baselines.
-
DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation
A diffusion model fuses per-person masks with pose keypoints to generate identity-preserving, two-person interactive videos from a single reference image, outperforming prior single-person-animation pipelines.
-
FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios
FlexiAct transfers actions from a reference video to an arbitrary target image, allowing changes in layout, skeleton, and viewpoint while keeping the target subject's appearance.
-
A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation
PAHA improves audio-driven avatar video generation by re-weighting training loss toward hands and face and by adding audio-video consistency classifiers during inference.
-
Satellite to GroundScape -- Large-scale Consistent Ground View Generation from Satellite Views
Satellite-to-ground generation is extended from single images to consistent multi-view sequences by adding satellite-guided and satellite-temporal conditioning to a frozen latent diffusion model, with a new 100k-pair dataset.
-
RealisDance-DiT: Simple yet Strong Baseline towards Controllable Character Animation in the Wild
Simple conditioning patches and training tricks on the Wan-2.1 model outperform specialized reference-network methods for controllable character animation, according to the paper's benchmarks.
-
AnyCharV: Bootstrap Controllable Character Video Generation with Fine-to-Coarse Guidance
AnyCharV is a two-stage diffusion method that places a reference character into a target video scene using pose and mask guidance, with a self-boosting stage that improves identity preservation.
-
FlexControl: Computation-Aware ControlNet with Differentiable Router for Text-to-Image Generation
FlexControl learns per-timestep, per-block gating for ControlNet control branches, using a FLOPs budget loss to cut compute while preserving or improving image fidelity.
-
Unpaired Deblurring via Decoupled Diffusion Model
A diffusion model that decouples structural features from blur patterns using unpaired target-domain images can deblur photos in unseen domains without paired training data.
-
EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion
EchoVideo preserves identity in generated human videos by pre-fusing face, image, and text features, then training with stochastic shallow-feature dropout to reduce copy-paste artifacts.
-
Video Depth Anything: Consistent Depth Estimation for Super-Long Videos
Video Depth Anything adapts Depth Anything V2 to produce temporally consistent depth for arbitrarily long videos using a temporal attention head, an optical-flow-free gradient loss, and key-frame-based stitching.
-
Ingredients: Blending Custom Photos with Video Diffusion Transformers
Ingredients adds a mask-supervised identity router to a video diffusion transformer, enabling multi-person videos from a few reference photos without per-identity fine-tuning.
-
Generative Inbetweening through Frame-wise Conditions-Driven Video Generation
FCVG generates stable inbetween frames by injecting linearly interpolated line-match and pose conditions into every denoising step of a pretrained video diffusion model.
-
OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation
A drag-style motion control method for 360 degree image-to-video generation, built on spherical trajectory estimation and joint fine-tuning of a pretrained video diffusion model.
-
Appearance Matching Adapter for Exemplar-based Semantic Image Synthesis in-the-Wild
AM-Adapter injects segmentation-derived matching costs into diffusion self-attention to transfer local object appearances from an exemplar to a target segmentation-driven scene.
-
FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation
A 2D-keypoint-conditioned diffusion model trained on a new 10M-image hand dataset enables controllable hand reposing, appearance transfer, novel view synthesis, and zero-shot hand video generation.
-
ControlFace: Harnessing Facial Parametric Control for Face Rigging
ControlFace performs zero-shot face rigging from 3DMM renderings using a dual-branch U-Net, a control mixer module, and reference control guidance, and reports the best average DECA re-inference error on FFHQ baselines.
-
DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses
A two-stage diffusion pipeline that enriches 2D pose guidance with generated depth and normal maps to achieve state-of-the-art human image animation.
-
StableAnimator: High-Quality Identity-Preserving Human Image Animation
StableAnimator uses a video diffusion model with face-embedding adapters and per-step latent optimization to generate pose-driven videos that preserve the reference person's identity end-to-end.
-
Time Step Generating: A Universal Synthesized Deepfake Image Detector
TSG classifies real versus synthetic images by feeding the image at a fixed noise timestep through a frozen diffusion U-Net and classifying its predicted noise map.
-
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation
EgoVid-5M is a 5M-clip curated egocentric video dataset with action and kinematic annotations, and EgoDreamer generates egocentric videos from text and camera control signals.
-
NanoControl: A Lightweight Framework for Precise and Efficient Control in Diffusion Transformer
NanoControl injects condition-specific key-value pairs into every attention block of Flux via a LoRA-style branch, claiming state-of-the-art controllability at 0.024% extra parameters and 0.029% extra FLOPs.
-
Animate-X++: Universal Character Image Animation with Dynamic Backgrounds
Animate-X++ turns cartoon images into pose-driven animations with text-controlled moving backgrounds, claiming state-of-the-art results on a new synthetic anthropomorphic benchmark.
-
StableAnimator++: Overcoming Pose Misalignment and Face Distortion for Human Image Animation
StableAnimator++ combines learnable SVD-guided pose alignment, a distribution-aware ID Adapter, and an HJB-based inference-time face optimizer to preserve identity in human image animation under severe pose misalignment.
-
LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video Diffusion
LiON-LoRA adds a learned scaling token to video-diffusion LoRA adapters, enabling linear and independent control of camera trajectory and object motion strength.
-
FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases
FixTalk adds two modules to a real-time GAN talking-head model, decoupling identity from motion to stop identity leakage while using a memory to recover details and reduce artifacts.
-
FramePrompt: In-context Controllable Animation with Zero Structural Changes
FramePrompt turns character animation into a video-continuation task by concatenating reference image, skeleton frames, and target frames into one sequence, then training the pretrained Wan-I2V model to generate only ...
-
FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
FullDiT2 accelerates FullDiT-style in-context conditioning for video by dynamic token selection and selective context caching, cutting per-step time by 2-3x with minimal quality loss.
-
Dual Prompting Image Restoration with Diffusion Transformers
DPIR combines lightweight conditioning with global-local CLIP visual prompts and T5 text in an SD3 diffusion transformer to improve image restoration quality.
-
RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control
Metric-scale depth alignment plus scene-constrained noise shaping improves camera controllability and video quality for image-to-video generation on RealEstate10K.
-
MetaFE-DE: Learning Meta Feature Embedding for Depth Estimation from Monocular Endoscopic Images
A temporal diffusion pretraining stage aligned with frame latents improves self-supervised monocular depth estimation in endoscopic video.
-
Motion Prompting: Controlling Video Generation with Motion Trajectories
A single-stage ControlNet on the Lumiere video model, conditioned only on dense point tracks, generalizes to sparse and dense trajectory control for object, camera, and transferred motions.
-
AccDiffusion v2: Towards More Accurate Higher-Resolution Diffusion Extrapolation
The paper reports a training-free extrapolation method that combines per-patch prompt selection from cross-attention masks, ControlNet edge guidance, and permuted dilated sampling to reduce repetition and distortion a...
-
CPA: Camera-pose-awareness Diffusion Transformer for Video Generation
CPA adds camera pose control to OpenSora-style diffusion video models by encoding a sparse Plücker motion field into a latent and injecting it into temporal attention layers.
-
Video Diffusion Transformers are In-Context Learners
Concatenating multiple video clips into one input and fine-tuning a LoRA adapter lets a pretrained video diffusion transformer produce consistent multi-scene videos from a single prompt.
-
Identity-Preserving Pose-Guided Character Animation via Facial Landmarks Transformation
FLT improves identity preservation in landmark-conditioned character animation by combining reference face shape with driving expressions in 3D and re-rendering landmarks.
-
Seeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attention
A diffusion transformer with joint attention over all camera views and frames, a lightweight BEV controller, and a re-weighted object loss reports state-of-the-art multi-view driving video generation (FVD 37.8 on nuScenes).
-
Efficient Diffusion Models: A Survey
The paper organizes research on efficient diffusion models into a taxonomy spanning algorithms, systems, and frameworks, and provides a curated reference list.
-
Parameter-Efficient Fine-Tuning for Foundation Models
A survey that categorizes and summarizes parameter-efficient fine-tuning methods across large language, vision, and multimodal models.
-
Text-to-Image Synthesis: A Decade Survey
A decade-spanning survey categorizes over 440 text-to-image papers by architecture, research problem, dataset, and evaluation metric.
Discussion (0). Continue with ORCID to comment.