PatternGSL introduces a learnable specification language for sewing patterns that lets vision-language models reconstruct explicit, simulation-ready 3D garments from single images, backed by a new 300K paired dataset.
hub
International Conference on Medical image computing and computer-assisted intervention , pages=
25 Pith papers cite this work. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
fields
cs.CV 13 cs.LG 3 physics.geo-ph 2 astro-ph.GA 1 astro-ph.IM 1 cs.AI 1 eess.IV 1 math.OC 1 physics.plasm-ph 1 stat.ML 1years
2026 25roles
background 1polarities
background 1representative citing papers
Linear-DPO replaces sigmoid utility with linear utility and adds EMA reference to improve preference alignment in diffusion and flow-matching text-to-image models.
A hypernetwork maps style motion embeddings to LoRA updates that stylize text-driven motion diffusion models with improved generalization to unseen styles via contrastive structuring of the style space.
A generative transfer framework using iterative path-wise tilting integrated with conditional flow matching recovers target entropic optimal transport couplings from reference samples, achieving O(δ) convergence in Wasserstein-1 distance.
Single-shot HDR is achieved by conditioning a video diffusion model on an LDR input to generate an exposure bracket and fusing the bracket with per-pixel weights from a lightweight UNet.
PODiff performs conditional diffusion in a fixed, variance-ordered POD latent space to enable efficient probabilistic super-resolution of high-dimensional scientific fields with lower memory and better-calibrated uncertainty than pixel-space or dropout baselines.
RetinaDiff uses phase correlation to create a motion-corrected physics prior and then applies a conditional diffusion model to reconstruct high-quality temporal LSCI from extremely limited frames.
PRISM lets pre-trained text-to-image models handle long prompts by breaking them into compositional parts, predicting noise separately, and merging outputs via energy-based conjunction, matching fine-tuned models while generalizing better to prompts over 500 tokens.
End-to-end framework reconstructs 4D whole-heart meshes from cine MRI using differentiable contour rendering and multi-scale temporal modeling, reporting 1.68 mm MAE and improved motion smoothness over prior methods.
A CNN trained on AREPO simulations and synthetic observations reverts edge-on 13CO spectral data to top-down views of the CMZ as a proof-of-concept for supervised reversion.
DiT-Pruning squares the weight term in the pruning saliency metric and adapts per-layer versus per-channel granularity, claiming strong quality retention for diffusion transformers at high sparsity.
AMUSE stabilizes Muon with time-varying schedule-free gradient evaluation, improving the performance-iteration Pareto frontier without learning-rate schedules.
A frequency-enhanced Vision Transformer with FDSA, FGMLP, WAFF, and FCSB modules delivers superior volumetric medical image segmentation performance and efficiency over prior state-of-the-art methods.
Observational and counterfactual distributions are linked by identical support and invariant features, enabling a flow-matching estimator with semiparametric efficiency correction to generate debiased counterfactuals from observations.
Adapting vision foundation models with LoRA and kurtosis-guided unsupervised test-time adaptation matches or exceeds domain-specific models for seismic denoising across multiple sites and unseen data.
Authors generated and released 3,000 unlabeled field and 4,000 labeled synthetic seismic datasets for global shelf-edge clinothems to enable deep learning for automated seismic stratigraphic interpretation.
A framework using patient-specific GMM normalization and uncertainty-gated anatomical attention for AAA thrombus segmentation reports SOTA in-distribution performance and substantially better multi-center generalization.
Semi-LAR is a semi-supervised contrastive learning framework with linear attention for nighttime flare removal that refines pseudo-labels via quality assessment and uses flare-aware patch-level contrastive losses.
Key vectors of concept text tokens in multimodal diffusion transformers encode an 'omission signal'; amplifying the omission-minus-presence direction during early denoising reduces object omission on FLUX.1-Dev and SD3.5-Medium.
ERPPO adds a DSA-based ambiguity estimator to MAPPO and switches between L1 and L2 entropy regularization to improve exploration and stability in non-stationary multi-dimensional observations.
A generative framework using geometric diffusion for brain networks and tabular diffusion for other organs integrates ICD-coded SDoH proxies to improve disease reasoning on UK Biobank data.
Neo, a cGAN, super-resolves HSC images to HST-like quality and improves galaxy morphological parameter accuracy by factors of 2-10.
Two deep learning autoregressive models predict the evolution of 2D ideal MHD instabilities while preserving key physical invariants such as global conservation trends and Alfvénic fluctuations.
citing papers explorer
-
PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D Garments
PatternGSL introduces a learnable specification language for sewing patterns that lets vision-language models reconstruct explicit, simulation-ready 3D garments from single images, backed by a new 300K paired dataset.
-
Linear-DPO: Linear Direct Preference Optimization for Diffusion and Flow-Matching Generative Models
Linear-DPO replaces sigmoid utility with linear utility and adds EMA reference to improve preference alignment in diffusion and flow-matching text-to-image models.
-
Stylized Text-to-Motion Generation via Hypernetwork-Driven Low-Rank Adaptation
A hypernetwork maps style motion embeddings to LoRA updates that stylize text-driven motion diffusion models with improved generalization to unseen styles via contrastive structuring of the style space.
-
Generative Transfer for Entropic Optimal Transport with Unknown Costs
A generative transfer framework using iterative path-wise tilting integrated with conditional flow matching recovers target entropic optimal transport couplings from reference samples, achieving O(δ) convergence in Wasserstein-1 distance.
-
Single-Shot HDR Recovery via a Video Diffusion Prior
Single-shot HDR is achieved by conditioning a video diffusion model on an LDR input to generate an exposure bracket and fusing the bracket with per-pixel weights from a lightweight UNet.
-
PODiff: Latent Diffusion in Proper Orthogonal Decomposition Space for Scientific Super-Resolution
PODiff performs conditional diffusion in a fixed, variance-ordered POD latent space to enable efficient probabilistic super-resolution of high-dimensional scientific fields with lower memory and better-calibrated uncertainty than pixel-space or dropout baselines.
-
Physics-Informed Conditional Diffusion for Motion-Robust Retinal Temporal Laser Speckle Contrast Imaging
RetinaDiff uses phase correlation to create a motion-corrected physics prior and then applies a conditional diffusion model to reconstruct high-quality temporal LSCI from extremely limited frames.
-
Long-Text-to-Image Generation via Compositional Prompt Decomposition
PRISM lets pre-trained text-to-image models handle long prompts by breaking them into compositional parts, predicting noise separately, and merging outputs via energy-based conjunction, matching fine-tuned models while generalizing better to prompts over 500 tokens.
-
Personalized 4D Whole-Heart Mesh Reconstruction from Cine MRI via Multi-Scale Temporal Modeling and Differentiable Contour Rendering
End-to-end framework reconstructs 4D whole-heart meshes from cine MRI using differentiable contour rendering and multi-scale temporal modeling, reporting 1.68 mm MAE and improved motion smoothness over prior methods.
-
IRIS: Deciphering Spectral-Line Imagery of the Galactic Center by Machine-Learning on Simulations
A CNN trained on AREPO simulations and synthetic observations reverts edge-on 13CO spectral data to top-down views of the CMZ as a proof-of-concept for supervised reversion.
-
Post-Training Pruning for Diffusion Transformers
DiT-Pruning squares the weight term in the pruning saliency metric and adapts per-layer versus per-channel granularity, claiming strong quality retention for diffusion transformers at high sparsity.
-
AMUSE: Anytime Muon with Stable Gradient Evaluation
AMUSE stabilizes Muon with time-varying schedule-free gradient evaluation, improving the performance-iteration Pareto frontier without learning-rate schedules.
-
FEFormer: Frequency-enhanced Vision Transformer for Generic Knowledge Extraction and Adaptive Feature Fusion in Volumetric Medical Image Segmentation
A frequency-enhanced Vision Transformer with FDSA, FGMLP, WAFF, and FCSB modules delivers superior volumetric medical image segmentation performance and efficiency over prior state-of-the-art methods.
-
Debiased Counterfactual Generation via Flow Matching from Observations
Observational and counterfactual distributions are linked by identical support and invariant features, enabling a flow-matching estimator with semiparametric efficiency correction to generate debiased counterfactuals from observations.
-
Parameter-Efficient Adaptation of Pre-Trained Vision Foundation Models for Active and Passive Seismic Data Denoising
Adapting vision foundation models with LoRA and kurtosis-guided unsupervised test-time adaptation matches or exceeds domain-specific models for seismic denoising across multiple sites and unseen data.
-
Massive-scale unlabeled field and labeled synthetic seismic datasets of global shelf-edge clinothems
Authors generated and released 3,000 unlabeled field and 4,000 labeled synthetic seismic datasets for global shelf-edge clinothems to enable deep learning for automated seismic stratigraphic interpretation.
-
Trust the Prior (or Not): Uncertainty-Aware Abdominal Aortic Aneurysm Segmentation
A framework using patient-specific GMM normalization and uncertainty-gated anatomical attention for AAA thrombus segmentation reports SOTA in-distribution performance and substantially better multi-center generalization.
-
Semi-LAR: Semi-supervised Contrastive Learning with Linear Attention for Removal of Nighttime Flares
Semi-LAR is a semi-supervised contrastive learning framework with linear attention for nighttime flare removal that refines pseudo-labels via quality assessment and uses flare-aware patch-level contrastive losses.
-
Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
Key vectors of concept text tokens in multimodal diffusion transformers encode an 'omission signal'; amplifying the omission-minus-presence direction during early denoising reduces object omission on FLUX.1-Dev and SD3.5-Medium.
-
ERPPO: Entropy Regularization-based Proximal Policy Optimization
ERPPO adds a DSA-based ambiguity estimator to MAPPO and switches between L1 and L2 entropy regularization to improve exploration and stability in non-stationary multi-dimensional observations.
-
Marrying Generative Model of Healthcare Events with Digital Twin of Social Determinants of Health for Disease Reasoning
A generative framework using geometric diffusion for brain networks and tabular diffusion for other organs integrates ICD-coded SDoH proxies to improve disease reasoning on UK Biobank data.
-
Photometric Super-Resolution for Improving Galaxy Morphological Measurements using Conditional Generative Adversarial Networks
Neo, a cGAN, super-resolves HSC images to HST-like quality and improves galaxy morphological parameter accuracy by factors of 2-10.
-
Autoregressive prediction of 2D MHD dynamics inferred from deep learning modeling
Two deep learning autoregressive models predict the evolution of 2D ideal MHD instabilities while preserving key physical invariants such as global conservation trends and Alfvénic fluctuations.
- Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous Reasoning
- Unifying Deep Stochastic Processes for Image Enhancement