ASTRA reframes transition-state search as guided diffusion inference that samples the isodensity surface between metastable basins and converges to first-order saddles via score differences and physical forces.
hub Mixed citations
Advances in neural information processing systems33, 6840–6851 (2020)
Mixed citation behavior. Most common role is background (62%).
hub tools
citation-role summary
citation-polarity summary
representative citing papers
ISPA reduces KV cache size by up to 50% in AR video models by transitioning layers to local attention and applying instance-specific least-squares weight modulation to compensate for lost history.
Introduces a Bridge latent interface that maps mismatched student latents into teacher space, enabling distillation from modern diffusion teachers to compact one-step students and raising SD 1.5 HPSv3 from 5.4 to 9.4 while keeping one-step speed.
DiSI disentangles stochastic interpolants into separate generation and regression paths, allowing controllable transitions between regression and generative image restoration with a unified few-step sampler.
EPC-3D-Diff synthesizes CT from CBCT via 3D latent diffusion with a physics-consistent projection equivariance constraint, reporting PSNR gains of +7.4 dB on phantom and +1.8 dB on clinical data.
SeamCam quantifies camouflage by computing one minus the highest IoU recoverable from category-conditioned detection proposals against a ground-truth mask, achieving 78.82% agreement with human judgments.
GDMD replaces raw-sample rewards with distillation-gradient rewards in RL-guided diffusion distillation, yielding 4-step models that surpass their multi-step teachers on GenEval and human preference metrics.
ViPS learns a universal, controllable pose space for auto-rigged meshes by transferring motion priors from video diffusion models, matching SOTA performance on plausibility and diversity while enabling zero-shot generalization.
A sequential diffusion framework generates controllable abdominal anatomies with a Volume Control Scalar that decouples organ size from body habitus, achieving Dice scores around 0.83 and reducing distributional mismatch by 73.6% in a hepatomegaly example.
Video diffusion models can be adapted into permutation-invariant generators for sparse novel view synthesis by treating the problem as video completion and removing temporal order cues.
Grounded Forcing introduces dual memory caching, reference-based positional embeddings, and proximity-weighted recaching to bridge stable semantics with local dynamics, improving long-range consistency in autoregressive video synthesis.
HiPolicy is a new hierarchical multi-frequency action chunking method for imitation learning that jointly generates coarse and fine action sequences with entropy-guided execution to improve performance and efficiency in robotic manipulation.
Generative dictionary retrieval decodes unseen Oracle Bone Script characters at 54.3% Top-10 accuracy by synthesizing plausible variants guided by character evolution principles.
WCog-VLA couples Game-CoT semantic reasoning with an aligned decoupled diffusion transformer to generate joint multi-agent trajectories and reaches 92.9 PDMS on NAVSIM.
Hybrid T2I generation with teacher-student pseudo-labeling plus VRAIN context-aware I2I rare-class editing improves LVIS instance segmentation AP, especially on rare categories.
GACR reformulates cloud removal as an observation-anchored residual inversion process with geo-contextual prior alignment to preserve semantic structures for improved downstream interpretation tasks.
MIC casts diffusion motion generation as stochastic control to support both objective-based and criterion-based constraints without training or differentiability requirements.
Adapting diffusion models causes hidden damage to unrelated concepts detectable via sparse autoencoders and zero-shot classification, and DriftScope provides a prompt-level token-drift diagnostic.
WaterGen decouples scene generation from medium degradation in a two-stage latent diffusion process to produce controllable realistic underwater images that improve downstream restoration and segmentation.
A lightweight transformer module learns quality-aware vectors from timestep and prompt embeddings to modulate adaptive LayerNorm in DiT blocks, yielding consistent image quality gains over baseline diffusion transformers.
DPDiff-AD conditions a diffusion model on local prototypes (via nearest aggregation) and global prototypes (via optimal transport) to model normality scalably in multi-class anomaly detection, reporting AUROC gains on 160-category data.
StreamEdit enables high-quality training-free video editing by adapting streaming video generation models with dual-branch fast sampling, self-attention bridge, cross-attention grounding, source-oriented guidance, and visual prompting, outperforming prior methods in few-step regimes.
DS-DiT decouples LR and Ref conditions in a Siamese diffusion transformer, adds patch-level weighting, and uses autoguidance to improve reference-based super-resolution for remote sensing images.
Pretrained autoencoders in medical latent diffusion encode discriminative features well for reconstruction but structure their latent spaces in ways that hinder classifier learning, a gap that persists across architectures and is not closed by domain fine-tuning.
citing papers explorer
-
A Priori Sampling of Transition States with Guided Diffusion
ASTRA reframes transition-state search as guided diffusion inference that samples the isodensity surface between metastable basins and converges to first-order saddles via score differences and physical forces.
-
Towards Memory-Efficient Autoregressive Video Generation via Instance-Specific Parametric Absorption
ISPA reduces KV cache size by up to 50% in AR video models by transitioning layers to local attention and applying instance-specific least-squares weight modulation to compensate for lost history.
-
Cross-Space Distillation: Teaching One-Step Students with Modern Diffusion Teachers
Introduces a Bridge latent interface that maps mismatched student latents into teacher space, enabling distillation from modern diffusion teachers to compact one-step students and raising SD 1.5 HPSv3 from 5.4 to 9.4 while keeping one-step speed.
-
Disentangling Generation and Regression in Stochastic Interpolants for Controllable Image Restoration
DiSI disentangles stochastic interpolants into separate generation and regression paths, allowing controllable transitions between regression and generative image restoration with a unified few-step sampler.
-
EPC-3D-Diff: Equivariant Physics Consistent Conditional 3D Latent Diffusion for CBCT to CT Synthesis
EPC-3D-Diff synthesizes CT from CBCT via 3D latent diffusion with a physics-consistent projection equivariance constraint, reporting PSNR gains of +7.4 dB on phantom and +1.8 dB on clinical data.
-
SeamCam: Quantifying Seamless Camouflage via Multi-Cue Visual Detectability
SeamCam quantifies camouflage by computing one minus the highest IoU recoverable from category-conditioned detection proposals against a ground-truth mask, achieving 78.82% agreement with human judgments.
-
Guiding Distribution Matching Distillation with Gradient-Based Reinforcement Learning
GDMD replaces raw-sample rewards with distillation-gradient rewards in RL-guided diffusion distillation, yielding 4-step models that surpass their multi-step teachers on GenEval and human preference metrics.
-
ViPS: Video-informed Pose Spaces for Auto-Rigged Meshes
ViPS learns a universal, controllable pose space for auto-rigged meshes by transferring motion priors from video diffusion models, matching SOTA performance on plausibility and diversity while enabling zero-shot generalization.
-
AbdomenGen: Sequential Volume-Conditioned Diffusion Framework for Abdominal Anatomy Generation
A sequential diffusion framework generates controllable abdominal anatomies with a Volume Control Scalar that decouples organ size from body habitus, achieving Dice scores around 0.83 and reducing distributional mismatch by 73.6% in a hepatomegaly example.
-
Novel View Synthesis as Video Completion
Video diffusion models can be adapted into permutation-invariant generators for sparse novel view synthesis by treating the problem as video completion and removing temporal order cues.
-
Grounded Forcing: Bridging Time-Independent Semantics and Proximal Dynamics in Autoregressive Video Synthesis
Grounded Forcing introduces dual memory caching, reference-based positional embeddings, and proximity-weighted recaching to bridge stable semantics with local dynamics, improving long-range consistency in autoregressive video synthesis.
-
HiPolicy: Hierarchical Multi-Frequency Action Chunking for Policy Learning
HiPolicy is a new hierarchical multi-frequency action chunking method for imitation learning that jointly generates coarse and fine action sequences with entropy-guided execution to improve performance and efficiency in robotic manipulation.
-
Decoding Ancient Oracle Bone Script via Generative Dictionary Retrieval
Generative dictionary retrieval decodes unseen Oracle Bone Script characters at 54.3% Top-10 accuracy by synthesizing plausible variants guided by character evolution principles.
-
WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving
WCog-VLA couples Game-CoT semantic reasoning with an aligned decoupled diffusion transformer to generate joint multi-agent trajectories and reaches 92.9 PDMS on NAVSIM.
-
TMI: Text-to-Image Meets Image-to-Image for Complementary Data Synthesis to Boost Long-Tailed Instance Segmentation
Hybrid T2I generation with teacher-student pseudo-labeling plus VRAIN context-aware I2I rare-class editing improves LVIS instance segmentation AP, especially on rare categories.
-
Interpretation-Oriented Cloud Removal via Observation-Anchored Residual Flow with Geo-Contextual Alignment
GACR reformulates cloud removal as an observation-anchored residual inversion process with geo-contextual prior alignment to preserve semantic structures for improved downstream interpretation tasks.
-
Training-free Controllable Human Motion Generation under Heterogeneous Constraints
MIC casts diffusion motion generation as stochastic control to support both objective-based and criterion-based constraints without training or differentiability requirements.
-
DriftScope: Measuring The Hidden Effects of Diffusion Model Adaptation
Adapting diffusion models causes hidden damage to unrelated concepts detectable via sparse autoencoders and zero-shot classification, and DriftScope provides a prompt-level token-drift diagnostic.
-
WaterGen: Decoupling Scene and Medium in Underwater Image Generation
WaterGen decouples scene generation from medium degradation in a two-stage latent diffusion process to produce controllable realistic underwater images that improve downstream restoration and segmentation.
-
Quality-Aware Modulation for Diffusion Transformers
A lightweight transformer module learns quality-aware vectors from timestep and prompt embeddings to modulate adaptive LayerNorm in DiT blocks, yielding consistent image quality gains over baseline diffusion transformers.
-
Dual Prototype-Conditioned Diffusion Model for Scalable Multi-Class Unsupervised Anomaly Detection in Large Category Spaces
DPDiff-AD conditions a diffusion model on local prototypes (via nearest aggregation) and global prototypes (via optimal transport) to model normality scalably in multi-class anomaly detection, reporting AUROC gains on 160-category data.
-
StreamEdit: Training-Free Video Editing via Few-Step Streaming Video Generation
StreamEdit enables high-quality training-free video editing by adapting streaming video generation models with dual-branch fast sampling, self-attention bridge, cross-attention grounding, source-oriented guidance, and visual prompting, outperforming prior methods in few-step regimes.
-
Learning to Balance: Decoupled Siamese Diffusion Transformer for Reference-Based Remote Sensing Image Super-Resolution
DS-DiT decouples LR and Ref conditions in a Siamese diffusion transformer, adds patch-level weighting, and uses autoguidance to improve reference-based super-resolution for remote sensing images.
-
The Learnability Gap in Medical Latent Diffusion
Pretrained autoencoders in medical latent diffusion encode discriminative features well for reconstruction but structure their latent spaces in ways that hinder classifier learning, a gap that persists across architectures and is not closed by domain fine-tuning.
-
LPH-VTON: Resolving the Structure-Texture Dilemma of Virtual Try-On via Latent Process Handover
LPH-VTON uses a single denoising process with staged handover from structure-biased to texture-biased diffusion models to improve both geometric alignment and textural fidelity in virtual try-on.
-
Head Forcing: Long Autoregressive Video Generation via Head Heterogeneity
Head Forcing assigns tailored KV cache strategies to local, anchor, and memory attention heads plus head-wise RoPE re-encoding to extend autoregressive video generation from seconds to minutes without training.
-
TOC-SR: Task-Optimal Compact diffusion for Image Super Resolution
TOC-SR builds a compact one-step diffusion model for image super-resolution achieving 6.6x fewer parameters and 2.8x fewer GMACs while maintaining strong reconstruction quality.
-
Diffusion Model as a Generalist Segmentation Learner
DiGSeg repurposes diffusion U-Nets as generalist segmentation learners by conditioning on image-mask latents and multi-scale CLIP text features, achieving strong cross-domain performance.
-
Geometry Preserving Loss Functions Promote Improved Adaptation of Blackbox Generative Model
Geometry-preserving losses based on tangent-space distances improve blackbox GAN adaptation to shifted distributions compared with standard losses.
-
Optimizing Diffusion Priors in Image Reconstruction from a Single Observation
Combining diffusion priors as a product-of-experts and optimizing exponents via Bayesian evidence maximization enables prior tuning from one observation in inverse imaging problems.
-
Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation
A decoupled memory branch with hybrid cues, cross-attention, and gating improves spatial consistency and data efficiency in long-horizon camera-trajectory video generation.
-
Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes
Empty-prompt (unconditional) inversion into a native text-to-3D model avoids "sink traps" and reconstructs and edits out-of-distribution 3D shapes more faithfully than text-guided inversion.
-
ReplicateAnyScene: Zero-Shot Video-to-3D Composition via Textual-Visual-Spatial Alignment
ReplicateAnyScene performs fully automated zero-shot video-to-compositional-3D reconstruction by cascading alignments of generic priors from vision foundation models across textual, visual, and spatial dimensions.
-
GIF: A Conditional Multimodal Generative Framework for IR Drop Imaging in Chip Layouts
GIF fuses geometrical image features and logical graph topology in a conditional diffusion model to generate high-quality IR drop images for chip layouts, outperforming prior ML methods on CircuitNet-N28 with SSIM 0.78, Pearson 0.95, PSNR 21.77, and NMAE 0.026.
-
Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch
Proprioception and multi-contact touch, fused with a physics-guided conditional diffusion model over a Structure-VAE SDF latent space, improve metric amodal object reconstruction under severe hand occlusion versus vision-only baselines.
-
What Matters in Virtual Try-Off? Dual-UNet Diffusion Model For Garment Reconstruction
A Dual-UNet diffusion model for virtual garment reconstruction from clothed images sets new benchmarks on VITON-HD and DressCode by optimizing Stable Diffusion variants, mask conditioning, and auxiliary losses.
-
GroundingAnomaly: Spatially-Grounded Diffusion for Few-Shot Anomaly Synthesis
GroundingAnomaly uses a Spatial Conditioning Module and Gated Self-Attention in a frozen diffusion U-Net to synthesize spatially accurate few-shot anomalies, reaching SOTA on MVTec AD and VisA for detection, segmentation, and instance detection.
-
Energy-based Tissue Manifolds for Longitudinal Multiparametric MRI Analysis
A baseline-trained energy manifold in MRI intensity space showed progressive drift toward the tumour regime in a recurrence case but not in a stable case, suggesting a segmentation-free longitudinal monitoring signal.
-
Beyond Semantics: Uncovering the Physics of Fakes via Universal Physical Descriptors for Cross-Modal Synthetic Detection
Five universal physical descriptors including Laplacian variance, Sobel statistics, and residual noise variance, when integrated as text encodings with CLIP, achieve up to 99.8% accuracy detecting synthetic images across GAN and diffusion model datasets.
-
Teaching an Agent to Sketch One Part at a Time
A multi-modal LM agent is trained to produce vector sketches part-by-part via supervised fine-tuning and process-reward RL on the new ControlSketch-Part dataset with automatic part annotations.
-
ComplexMimic: Human-Scene Interaction Imitation in Complex 3D Environments
Dual-expert RL plus difficulty-aware multi-teacher distillation improves physics-based human–scene interaction imitation under complex 3D geometry versus prior single-policy baselines.
-
MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts
MEPA adds token-routed MoE and residual self-supervised feature alignment to VAR models, reporting better FID on ImageNet 256x256 with half the training epochs and fewer parameters than dense baselines.
-
DVG-WM: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation
DVG-WM disentangles dynamics learning from visual synthesis via flow matching and latent degradation to deliver faster, higher-quality video predictions for robotic manipulation.
-
Stochastic Optimal Control Sampling for Diffusion Inverse Problems
SOCS derives per-step closed-form control signals from stochastic optimal control to steer diffusion sampling trajectories toward measurements while preserving the generative prior.
-
ULF-Synth: Physics-Guided Ultra-Low-Field MRI Enhancement for Pediatric Neuroimaging
ULF-Synth creates synthetic ULF images from HF volumes and uses a frequency-domain loss to train models that generalize to real 64mT ULF scans, boosting segmentation and radiologist ratings.
-
Structural Energy Guidance for View-Consistent Text-to-3D Generation
SEGS constructs structural energy in the PCA subspace of U-Net features and injects its gradient into the denoising process to improve multi-view consistency in text-to-3D generation.
-
AnimeAdapter: A Modular Adapter for Appearance-Consistent Anime Character Generation
AnimeAdapter is a modular adapter for Stable Diffusion that enables appearance-consistent anime character generation from a single reference image using semantic-selective local attention and pose-aware conditioning, plus a new Danbooru-derived dataset.
-
X-Imitator: Spatial-Aware Imitation Learning via Bidirectional Action-Pose Interaction
X-Imitator is a bidirectional action-pose interaction framework for spatial-aware imitation learning that outperforms vanilla policies and explicit pose guidance on 24 simulated and 3 real-world robotic tasks.
-
diffGHOST: Diffusion based Generative Hedged Oblivious Synthetic Trajectories
diffGHOST is a conditional diffusion model that segments learned latent space to identify and mitigate memorization of critical trajectory samples, aiming to deliver privacy guarantees alongside data utility.
-
DepthPilot: From Controllability to Interpretability in Colonoscopy Video Generation
DepthPilot generates physically consistent and clinically interpretable colonoscopy videos by injecting depth priors into diffusion models through parameter-efficient fine-tuning and replacing linear denoising weights with adaptive splines.