RefineAnything is a multimodal diffusion model using Focus-and-Refine crop-and-resize with blended paste-back to achieve high-fidelity local image refinement and near-perfect background preservation.
In: ICML (2024)
11 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CV 11years
2026 11representative citing papers
Pretrained VAE latents have far lower transmission–reflection coherence than pixels, so flow matching on FLUX plus cycle and contrastive losses can separate both layers in one pass.
FlowCIR frames ZS-CIR as conditional flow matching transport on fixed VLM embeddings plus an inference-time Multi-Negative Steering fix for negation, reporting competitive benchmark results at far lower training cost.
Anti-Prompt protects images from text-guided image-to-video generation by suppressing text-conditioned attention during denoising, producing visible generation failures.
DriveWeaver performs point-conditioned video inpainting with a global-to-local hierarchical strategy to insert controllable vehicles into autonomous driving simulations and extracts 3D Gaussians for real-time rendering.
A saliency-guided warp-unwarp method reallocates spatial representation to preserve fine structures in latent diffusion models for image-to-image translation.
Mural transfers knowledge from a frozen LLM to text-to-image synthesis via MoT shared attention, achieving 0.85 GenEval, 86.75 DPG-Bench, and 0.66 WISE while exhibiting emergent behaviors without multimodal or reasoning supervision.
LimeCross enables text-guided editing of individual layers in composite images by conditioning on cross-layer context via bi-stream attention while preserving layer integrity and introducing the LayerEditBench benchmark.
Refinement via Regeneration (RvR) reformulates image refinement in unified multimodal models as conditional regeneration using prompt and semantic tokens from the initial image, yielding higher alignment scores than editing-based methods.
LUNA is an LBS-free neural animation model that maps 2D controls to 3D Gaussian deformations via a transformer motion regressor and hybrid supervision for realistic motion and zero-shot generalization.
ALM integrates likelihood maximization and acceleration into diffusion reverse sampling to enable globally coherent generation from incomplete inputs.
citing papers explorer
-
RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details
RefineAnything is a multimodal diffusion model using Focus-and-Refine crop-and-resize with blended paste-back to achieve high-fidelity local image refinement and near-perfect background preservation.
-
PRISM: Latent Composition Consistency for Single-Image Reflection Removal
Pretrained VAE latents have far lower transmission–reflection coherence than pixels, so flow matching on FLUX plus cycle and contrastive losses can separate both layers in one pass.
-
FlowCIR: Semantic Transport via Flow Matching for Zero-Shot Composed Image Retrieval
FlowCIR frames ZS-CIR as conditional flow matching transport on fixed VLM embeddings plus an inference-time Multi-Negative Steering fix for negation, reporting competitive benchmark results at far lower training cost.
-
Anti-Prompt: Image Protection against Text-Guided Image-to-Video Generation
Anti-Prompt protects images from text-guided image-to-video generation by suppressing text-conditioned attention during denoising, producing visible generation failures.
-
DriveWeaver: Point-Conditioned Video Inpainting for Controllable Vehicle Insertion in Autonomous Driving Simulation
DriveWeaver performs point-conditioned video inpainting with a global-to-local hierarchical strategy to insert controllable vehicles into autonomous driving simulations and extracts 3D Gaussians for real-time rendering.
-
WarpI2I: Image Warping for Image-to-Image Translation
A saliency-guided warp-unwarp method reallocates spatial representation to preserve fine structures in latent diffusion models for image-to-image translation.
-
Mural: Transferring LLM knowledge to image generation via Mixture-of-Transformers
Mural transfers knowledge from a frozen LLM to text-to-image synthesis via MoT shared attention, achieving 0.85 GenEval, 86.75 DPG-Bench, and 0.66 WISE while exhibiting emergent behaviors without multimodal or reasoning supervision.
-
LimeCross: Context-Conditioned Layered Image Editing with Structural Consistency
LimeCross enables text-guided editing of individual layers in composite images by conditioning on cross-layer context via bi-stream attention while preserving layer integrity and introducing the LayerEditBench benchmark.
-
Refinement via Regeneration: Enlarging Modification Space Boosts Image Refinement in Unified Multimodal Models
Refinement via Regeneration (RvR) reformulates image refinement in unified multimodal models as conditional regeneration using prompt and semantic tokens from the initial image, yielding higher alignment scores than editing-based methods.
-
LUNA: Learning Universal 3D Human Animation Beyond Skinning
LUNA is an LBS-free neural animation model that maps 2D controls to 3D Gaussian deformations via a transformer motion regressor and hybrid supervision for realistic motion and zero-shot generalization.
-
Accelerated Likelihood Maximization for Diffusion-based Versatile Content Generation
ALM integrates likelihood maximization and acceleration into diffusion reverse sampling to enable globally coherent generation from incomplete inputs.