REVIEW 8 cited by
MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
A hallmark of human intelligence is the ability to create complex artifacts through structured multi-step processes. Generating procedural tutorials with AI is a longstanding but challenging goal, facing three key obstacles: (1) scarcity of multi-task procedural datasets, (2) maintaining logical continuity and visual consistency between steps, and (3) generalizing across multiple domains. To address these challenges, we propose a multi-domain dataset covering 21 tasks with over 24,000 procedural sequences. Building upon this foundation, we introduce MakeAnything, a framework based on the diffusion transformer (DIT), which leverages fine-tuning to activate the in-context capabilities of DIT for generating consistent procedural sequences. We introduce asymmetric low-rank adaptation (LoRA) for image generation, which balances generalization capabilities and task-specific performance by freezing encoder parameters while adaptively tuning decoder layers. Additionally, our ReCraft model enables image-to-process generation through spatiotemporal consistency constraints, allowing static images to be decomposed into plausible creation sequences. Extensive experiments demonstrate that MakeAnything surpasses existing methods, setting new performance benchmarks for procedural generation tasks.
Forward citations
Cited by 8 Pith papers
-
TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy
TryOnCrafter is the first DiT-based framework for camera-controllable video virtual try-on via a renderable 4D try-on proxy distilled from 2D priors into 3DGS avatar animated with SMPL-X.
-
TanDiT: Tangent-Plane Diffusion Transformer for High-Quality 360{\deg} Panorama Generation
TanDiT generates high-quality 360-degree panoramas by jointly generating grids of tangent-plane views with a single diffusion transformer and refining them with a pretrained model.
-
Autoregressive Images Watermarking through Lexical Biasing: An Approach Resistant to Regeneration Attack
LBW embeds watermarks into autoregressive image token maps by biasing token sampling toward a secret green list and detects them with a z-test on green-token counts.
-
DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers
DiffDecompose recovers foreground and background layers from alpha-composited images using in-context diffusion with position encoding cloning, trained and evaluated on a new six-task synthetic dataset.
-
OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data
OmniConsistency is a style-agnostic consistency module for Flux that preserves structure and details during stylization with arbitrary LoRAs, reaching GPT-4o-level content consistency.
-
WorldWander: Bridging Egocentric and Exocentric Worlds in Video Generation
A bidirectional egocentric-to-exocentric video translation framework trained with in-context attention on a new synthetic+real dataset, with evaluation flaws around reference leakage and missing direct baselines.
-
FeRA: Frequency-Energy Constrained Routing for Effective Diffusion Adaptation Fine-Tuning
A frequency-energy router that blends LoRA experts according to the latent's bandwise energy improves diffusion fine-tuning quality and style consistency across multiple backbones.
-
RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers
A decoupled-attention adapter transfers image-pair edits to new photos in diffusion transformers, trained with a new 218-task visual editing dataset.
Discussion (0). Sign in to comment.