REVIEW 16 cited by
UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Achieving machine autonomy and human control often represent divergent objectives in the design of interactive AI systems. Visual generative foundation models such as Stable Diffusion show promise in navigating these goals, especially when prompted with arbitrary languages. However, they often fall short in generating images with spatial, structural, or geometric controls. The integration of such controls, which can accommodate various visual conditions in a single unified model, remains an unaddressed challenge. In response, we introduce UniControl, a new generative foundation model that consolidates a wide array of controllable condition-to-image (C2I) tasks within a singular framework, while still allowing for arbitrary language prompts. UniControl enables pixel-level-precise image generation, where visual conditions primarily influence the generated structures and language prompts guide the style and context. To equip UniControl with the capacity to handle diverse visual conditions, we augment pretrained text-to-image diffusion models and introduce a task-aware HyperNet to modulate the diffusion models, enabling the adaptation to different C2I tasks simultaneously. Trained on nine unique C2I tasks, UniControl demonstrates impressive zero-shot generation abilities with unseen visual conditions. Experimental results show that UniControl often surpasses the performance of single-task-controlled methods of comparable model sizes. This control versatility positions UniControl as a significant advancement in the realm of controllable visual generation.
Forward citations
Cited by 16 Pith papers
-
UNITY: Attention Flow Networks for Adaptive Conditioning in Diffusion
UNITY is a two-stage adapter with Morphable Attention Flow networks for efficient single and composite conditioning in diffusion-based image generation.
-
Dual Recursive Feedback on Generation and Appearance Latents for Pose-Robust Text-to-Image Diffusion
Dual Recursive Feedback steers the shared noise of appearance and generation latents during sampling, improving structure-appearance fusion in training-free controllable text-to-image diffusion.
-
Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas
Canvas3D lets users arrange objects in a 3D canvas generated from a text prompt, then feeds depth, skeleton, and lighting constraints to diffusion models to produce images that match the layout.
-
See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
Fine-tuning LVLMs on original/altered image pairs with targeted visual instructions reduces object and attribute hallucinations on POPE, LLaVA-Bench, and MMHal-Bench.
-
AnyI2V: Animating Any Conditional Image with Motion Control
AnyI2V animates arbitrary conditional images with user-defined trajectories by injecting debiased diffusion features and aligning attention queries across frames, without training.
-
Heeding the Inner Voice: Aligning ControlNet Training via Intermediate Features Feedback
InnerControl trains lightweight probes on intermediate UNet features to enforce control alignment throughout the denoising trajectory, improving controllability for edges and depth.
-
Controllable Coupled Image Generation via Diffusion Models
A cross-attention control method that couples backgrounds across multiple generated images by blending LLM-extracted background and entity prompts with a time-varying weight optimized for background similarity and tex...
-
Image Editing As Programs with Diffusion Models
IEAP decomposes complex editing instructions into atomic operations executed sequentially on a diffusion transformer, and reports state-of-the-art results on MagicBrush and AnyEdit.
-
Minimal Impact ControlNet: Advancing Multi-ControlNet Integration
MIControlNet balances silent-region data, MGDA-style feature fusion, and a Jacobian-symmetry loss to reduce multi-ControlNet conflicts and improve multi-condition FID.
-
ImmunoDiff: A Diffusion Model for Immunotherapy Response Prediction in Lung Cancer
An anatomy- and clinical-conditioned diffusion model that synthesizes post-treatment CT and uses its features to improve immunotherapy response prediction in NSCLC.
-
Zero-to-Hero: Zero-Shot Initialization Empowering Reference-Based Video Appearance Editing
A reference-based video editing pipeline that guides cross-image attention with diffusion correspondence, then trains a per-video restoration model to clean up the zero-shot output.
-
FlowForm: Synergizing Fluid Physics with Topological Consistency for Satellite Flood Synthesis
FlowForm is a diffusion model that adds shallow-water-equation penalties and terrain-conditioned adapters to synthesize flood satellite images, and reports top scores on a new 10,000-pair dataset.
-
High-$Q$ superconducting resonators fabricated in an industry-scale semiconductor-fabrication facility
Superconducting niobium and tantalum resonators fabricated in a 200 mm semiconductor production line demonstrate Q-factors above 10^6 in the single-photon regime.
-
LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs
A multimodal-LLM planner, polygon-based layout masks, and SVD-based structure injection combine to improve spatial and textual control of pre-trained text-to-image diffusion models.
-
Ovis-U1 Technical Report
A 3B unified multimodal model with a diffusion decoder and bidirectional refiner achieves competitive understanding, generation, and editing benchmark scores.
-
Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis
A video diffusion model, HunyuanVideo-I2V, is adapted with mixup transitions, frame-skip position embeddings, and attention masking to outperform image-only models on several controllable image generation benchmarks.
Discussion (0). Sign in to comment.