REVIEW 8 cited by
Mixture of Diffusers for scene composition and high resolution image generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Diffusion methods have been proven to be very effective to generate images while conditioning on a text prompt. However, and although the quality of the generated images is unprecedented, these methods seem to struggle when trying to generate specific image compositions. In this paper we present Mixture of Diffusers, an algorithm that builds over existing diffusion models to provide a more detailed control over composition. By harmonizing several diffusion processes acting on different regions of a canvas, it allows generating larger images, where the location of each object and style is controlled by a separate diffusion process.
Forward citations
Cited by 8 Pith papers
-
Latent Swap Joint Diffusion for 2D Long-Form Latent Generation
A training-free latent swap method that replaces averaging with binary swapping in joint diffusion, improving long-form audio spectrum and panorama generation.
-
CASR: A Robust Cyclic Framework for Arbitrary Large-Scale Super-Resolution with Distribution Alignment and Self-Similarity Awareness
CASR enables stable arbitrary-scale super-resolution by breaking extreme magnifications into cyclic in-distribution transitions with SSAM for structural distribution alignment and SARM for texture self-similarity pres...
-
StableCodec: Taming One-Step Diffusion for Extreme Image Compression
A one-step diffusion codec that compresses noisy latents at 64x and decodes with a single denoising step, setting state-of-the-art FID, KID, and DISTS at ultra-low bitrates.
-
JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers
JointDiT is a Flux-based diffusion transformer that jointly generates RGB images and depth maps, and also handles depth estimation and depth-to-image generation by controlling the noise level of each branch.
-
DMM: Building a Versatile Image Generation Model via Distillation-Based Model Merging
DMM trains a single style-promptable diffusion model to reproduce the outputs of multiple teacher models, achieving a merged-model FIDt of 77.51 versus a reference of 74.91.
-
SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration
A 2.48B-parameter diffusion transformer with shifted-window attention and a causal video autoencoder reports competitive perceptual-quality video restoration across synthetic, real-world, and AI-generated benchmarks.
-
PiCo: Enhancing Text-Image Alignment with Improved Noise Selection and Precise Mask Control in Diffusion Models
PiCo improves text-to-image alignment by scoring and selecting initial noise seeds with CLIPSeg-based global and concept scores, then modulating cross-attention maps with CLIPSeg pixel masks and exclusive masks.
-
Observable Performance Does Not Fully Reflect Adaptive System Organization: A Multi-Level Analysis of Gait Dynamics Under Occlusal Constraint
In one Parkinson's patient, six occlusal probes produce overlapping gait scores and UMAP embeddings, so observable performance does not uniquely identify adaptive system state under VDO constraint.
Discussion (0). Continue with ORCID to comment.