COSY uses independent per-component 3DGS generators plus context tokens to achieve disentangled semantic editing of human heads without masks or classifiers.
Advances in neural in- formation processing systems27(2014)
9 Pith papers cite this work. Polarity classification is still indexing.
years
2026 9representative citing papers
GDMD replaces raw-sample rewards with distillation-gradient rewards in RL-guided diffusion distillation, yielding 4-step models that surpass their multi-step teachers on GenEval and human preference metrics.
InterTalk generates real-time multi-participant conversational talking-face videos by iteratively refining each person's lip, eye, and head motions from other participants' audio and motion feedback.
LPH-VTON uses a single denoising process with staged handover from structure-biased to texture-biased diffusion models to improve both geometric alignment and textural fidelity in virtual try-on.
A Dual-UNet diffusion model for virtual garment reconstruction from clothed images sets new benchmarks on VITON-HD and DressCode by optimizing Stable Diffusion variants, mask conditioning, and auxiliary losses.
GroundingAnomaly uses a Spatial Conditioning Module and Gated Self-Attention in a frozen diffusion U-Net to synthesize spatially accurate few-shot anomalies, reaching SOTA on MVTec AD and VisA for detection, segmentation, and instance detection.
ULF-Synth creates synthetic ULF images from HF volumes and uses a frequency-domain loss to train models that generalize to real 64mT ULF scans, boosting segmentation and radiologist ratings.
AlloSR² claims state-of-the-art one-step real-world super-resolution by SNR-guided trajectory init, velocity regularization (FATC), and allomorphic self-adversarial distillation that preserves flow-matching generative priors.
DiMSO trains a fully connected network to map Gaussian noise to a real tabular distribution, claiming state-of-the-art similarity with hundreds of times faster generation than CTGAN/TVAE.
citing papers explorer
-
COSY: Compositional 3DGS Synthesis for Disentangled Human Head Editing
COSY uses independent per-component 3DGS generators plus context tokens to achieve disentangled semantic editing of human heads without masks or classifiers.
-
Guiding Distribution Matching Distillation with Gradient-Based Reinforcement Learning
GDMD replaces raw-sample rewards with distillation-gradient rewards in RL-guided diffusion distillation, yielding 4-step models that surpass their multi-step teachers on GenEval and human preference metrics.
-
Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation
InterTalk generates real-time multi-participant conversational talking-face videos by iteratively refining each person's lip, eye, and head motions from other participants' audio and motion feedback.
-
LPH-VTON: Resolving the Structure-Texture Dilemma of Virtual Try-On via Latent Process Handover
LPH-VTON uses a single denoising process with staged handover from structure-biased to texture-biased diffusion models to improve both geometric alignment and textural fidelity in virtual try-on.
-
What Matters in Virtual Try-Off? Dual-UNet Diffusion Model For Garment Reconstruction
A Dual-UNet diffusion model for virtual garment reconstruction from clothed images sets new benchmarks on VITON-HD and DressCode by optimizing Stable Diffusion variants, mask conditioning, and auxiliary losses.
-
GroundingAnomaly: Spatially-Grounded Diffusion for Few-Shot Anomaly Synthesis
GroundingAnomaly uses a Spatial Conditioning Module and Gated Self-Attention in a frozen diffusion U-Net to synthesize spatially accurate few-shot anomalies, reaching SOTA on MVTec AD and VisA for detection, segmentation, and instance detection.
-
ULF-Synth: Physics-Guided Ultra-Low-Field MRI Enhancement for Pediatric Neuroimaging
ULF-Synth creates synthetic ULF images from HF volumes and uses a frequency-domain loss to train models that generalize to real 64mT ULF scans, boosting segmentation and radiologist ratings.
-
Allo{SR}$^2$: Rectifying One-Step Super-Resolution to Stay Real via Allomorphic Generative Flows
AlloSR² claims state-of-the-art one-step real-world super-resolution by SNR-guided trajectory init, velocity regularization (FATC), and allomorphic self-adversarial distillation that preserves flow-matching generative priors.
-
Synthesizing real-world distributions from high-dimensional Gaussian Noise with Fully Connected Neural Network
DiMSO trains a fully connected network to map Gaussian noise to a real tabular distribution, claiming state-of-the-art similarity with hundreds of times faster generation than CTGAN/TVAE.