MetaEarth-MM unifies multi-modal remote sensing image generation and any-to-any translation across five modalities via scene-centered joint modeling on the new EarthMM dataset.
Instructpix2pix: Learning to follow image editing instructions
6 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 6roles
background 1polarities
background 1representative citing papers
Defines SML task for localizing semantic edits and proposes TRACE framework with semantic anchoring, perturbation sensing, and constrained reasoning that outperforms prior IML methods on a custom benchmark.
SwiftAudio performs caption-only distillation of a one-step TTA diffusion model by adapting VSD to audio with temporal smoothness regularization, achieving SOTA among one-step methods on AudioCaps and Clotho using ~45K captions.
A 3D-warped synthetic triplet dataset plus GRPO post-training on real portraits yields state-of-the-art identity-preserving makeup transfer, evaluated on a new diverse BeautyBench benchmark.
StructDiff adds adaptive receptive fields and 3D positional encoding to a single-scale diffusion model to preserve structure and enable spatial control in single-image generation.
A hierarchical multi-agent framework that fuses three CLIP-based retrieval views and then applies tournament-style test-time reasoning reports the best published scores on CIRR, CIRCO, and FashionIQ.
citing papers explorer
-
MetaEarth-MM: Unified Multimodal Remote Sensing Image Generation with Scene-centered Joint Modeling
MetaEarth-MM unifies multi-modal remote sensing image generation and any-to-any translation across five modalities via scene-centered joint modeling on the new EarthMM dataset.
-
Semantic Manipulation Localization
Defines SML task for localizing semantic edits and proposes TRACE framework with semantic anchoring, perturbation sensing, and constrained reasoning that outperforms prior IML methods on a custom benchmark.
-
SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation
SwiftAudio performs caption-only distillation of a one-step TTA diffusion model by adapting VSD to audio with temporal smoothness regularization, achieving SOTA among one-step methods on AudioCaps and Clotho using ~45K captions.
-
From Synthetic to Real: Toward Identity-Consistent Makeup Transfer with Synthetic and Real Data
A 3D-warped synthetic triplet dataset plus GRPO post-training on real portraits yields state-of-the-art identity-preserving makeup transfer, evaluated on a new diverse BeautyBench benchmark.
-
StructDiff: A Structure-Preserving and Spatially Controllable Diffusion Model for Single-Image Generation
StructDiff adds adaptive receptive fields and 3D positional encoding to a single-scale diffusion model to preserve structure and enable spatial control in single-image generation.
-
DeliCIR: Memory-Guided Test-Time Deliberation via Multi-Agent Collaboration for Composed Image Retrieval
A hierarchical multi-agent framework that fuses three CLIP-based retrieval views and then applies tournament-style test-time reasoning reports the best published scores on CIRR, CIRCO, and FashionIQ.