MetaEarth-MM unifies multi-modal remote sensing image generation and any-to-any translation across five modalities via scene-centered joint modeling on the new EarthMM dataset.
Scaling rectified flow transformers for high-resolution image synthesis
10 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
representative citing papers
AEGIS localizes sparse semantic-injecting attention heads in diffusion models and applies similarity-aware repulsion at those heads to block visual synonym jailbreaks while preserving benign generation.
Hierarchical anti-aesthetic adversarial noise, guided by global and face-local preference reward models, degrades customized diffusion outputs and reduces facial identity leakage more than prior cloaking methods.
Hydra stabilizes multi-concept backdoor attacks in diffusion models via evolutionary trigger search in text encoder space and trigger-clean regularization during multi-task fine-tuning, achieving high attack success while preserving clean image quality.
D2-CDIG conditions diffusion models on DEM and cloud-fog priors to generate controlled remote sensing images with decoupled terrain and atmospheric control.
A 3D-warped synthetic triplet dataset plus GRPO post-training on real portraits yields state-of-the-art identity-preserving makeup transfer, evaluated on a new diverse BeautyBench benchmark.
ODP-Net uses instance-aware orthogonal decomposition, perturbation-based purification, and manifold alignment to separate universal forgery traces, generator fingerprints, and semantics, achieving SOTA on unseen architectures like Stable Diffusion 3.
MAFL uses adversarial training to suppress pattern and content biases, guiding models to learn shared generative features for better cross-model generalization in detecting AI images.
HardFlow turns hard constraint enforcement during flow-matching sampling into a tractable terminal-time trajectory optimization problem using optimal control.
A coarse-to-fine hybrid MMDiT/DiT audio editor trained with rectified flow matching improves fidelity and cuts edit time versus prior instruction-guided baselines on synthetic overlapping-event tasks.
citing papers explorer
-
MetaEarth-MM: Unified Multimodal Remote Sensing Image Generation with Scene-centered Joint Modeling
MetaEarth-MM unifies multi-modal remote sensing image generation and any-to-any translation across five modalities via scene-centered joint modeling on the new EarthMM dataset.
-
AEGIS: A Mechanism-Guided Defense against Visual Synonym Jailbreaks in Text-to-Image Models
AEGIS localizes sparse semantic-injecting attention heads in diffusion models and applies similarity-aware repulsion at those heads to block visual synonym jailbreaks while preserving benign generation.
-
Hierarchical Anti-Aesthetics: Protecting Facial Privacy against Customized Diffusion Models
Hierarchical anti-aesthetic adversarial noise, guided by global and face-local preference reward models, degrades customized diffusion outputs and reduces facial identity leakage more than prior cloaking methods.
-
Awakening the Hydra: Stabilizing Multi-Concept Backdoor Injection in Text-to-Image Diffusion Models
Hydra stabilizes multi-concept backdoor attacks in diffusion models via evolutionary trigger search in text encoder space and trigger-clean regularization during multi-task fine-tuning, achieving high attack success while preserving clean image quality.
-
D2-CDIG: Controlled Diffusion Remote Sensing Image Generation with Dual Priors of DEM and Cloud-Fog
D2-CDIG conditions diffusion models on DEM and cloud-fog priors to generate controlled remote sensing images with decoupled terrain and atmospheric control.
-
From Synthetic to Real: Toward Identity-Consistent Makeup Transfer with Synthetic and Real Data
A 3D-warped synthetic triplet dataset plus GRPO post-training on real portraits yields state-of-the-art identity-preserving makeup transfer, evaluated on a new diverse BeautyBench benchmark.
-
Decoupling Semantics and Fingerprints: A Universal Representation for AI-Generated Image Detection
ODP-Net uses instance-aware orthogonal decomposition, perturbation-based purification, and manifold alignment to separate universal forgery traces, generator fingerprints, and semantics, achieving SOTA on unseen architectures like Stable Diffusion 3.
-
Combating Pattern and Content Bias: Adversarial Feature Learning for Generalized AI-Generated Image Detection
MAFL uses adversarial training to suppress pattern and content biases, guiding models to learn shared generative features for better cross-model generalization in detecting AI images.
-
HardFlow: Hard-Constrained Sampling for Flow-Matching Models via Trajectory Optimization
HardFlow turns hard constraint enforcement during flow-matching sampling into a tractable terminal-time trajectory optimization problem using optimal control.
-
RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers
A coarse-to-fine hybrid MMDiT/DiT audio editor trained with rectified flow matching improves fidelity and cuts edit time versus prior instruction-guided baselines on synthetic overlapping-event tasks.