REVIEW 27 cited by
One-Step Image Translation with Text-to-Image Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
One-Step Image Translation with Text-to-Image Models
read the original abstract
In this work, we address two limitations of existing conditional diffusion models: their slow inference speed due to the iterative denoising process and their reliance on paired data for model fine-tuning. To tackle these issues, we introduce a general method for adapting a single-step diffusion model to new tasks and domains through adversarial learning objectives. Specifically, we consolidate various modules of the vanilla latent diffusion model into a single end-to-end generator network with small trainable weights, enhancing its ability to preserve the input image structure while reducing overfitting. We demonstrate that, for unpaired settings, our model CycleGAN-Turbo outperforms existing GAN-based and diffusion-based methods for various scene translation tasks, such as day-to-night conversion and adding/removing weather effects like fog, snow, and rain. We extend our method to paired settings, where our model pix2pix-Turbo is on par with recent works like Control-Net for Sketch2Photo and Edge2Image, but with a single-step inference. This work suggests that single-step diffusion models can serve as strong backbones for a range of GAN learning objectives. Our code and models are available at https://github.com/GaParmar/img2img-turbo.
Forward citations
Cited by 27 Pith papers
-
DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
DF3DV-1K supplies 1,048 scenes with clean and cluttered image pairs plus a challenging 41-scene subset to benchmark and improve distractor-free radiance field methods.
-
FieryGS: In-the-Wild Fire Synthesis with Physics-Integrated Gaussian Splatting
FieryGS integrates LLM-based material reasoning, volumetric combustion simulation, and a unified renderer with 3D Gaussian Splatting to generate physically plausible and user-controllable fire in in-the-wild scenes.
-
ProDiG: Progressive Diffusion-Guided Gaussian Splatting for Aerial to Ground Reconstruction
ProDiG progressively transforms aerial Gaussian splats into coherent ground-level 3D reconstructions via diffusion guidance and specialized attention modules.
-
ChopGrad: Pixel-Wise Losses for Latent Video Diffusion via Truncated Backpropagation
ChopGrad truncates backpropagation to local frame windows in video diffusion models, reducing memory from linear in frame count to constant while enabling pixel-wise loss fine-tuning.
-
LooseRoPE: Content-aware Attention Manipulation for Semantic Harmonization
LooseRoPE modulates RoPE in diffusion attention maps to continuously trade off between preserving a pasted object's identity and harmonizing it with its new surroundings.
-
WarpI2I: Image Warping for Image-to-Image Translation
A saliency-guided warp-unwarp method reallocates spatial representation to preserve fine structures in latent diffusion models for image-to-image translation.
-
Semantic-Aware, Physics-Informed, Geometry-Grounded Weather Video Synthesis
A new framework factorizes weather video synthesis into semantic appearance anchoring, physics-informed Gaussian particle simulation under gravity/wind/turbulence, and geometry-grounded alignment to produce diverse re...
-
Addressing Detail Bottlenecks in Latent Diffusion for RGB-to-SWIR Image Translation
Introduces SCAE with skip connections and LGE to fix detail loss in LDMs for RGB-to-SWIR translation, yielding up to 2x mAP gains and 3.4x on small objects while reaching SOTA FID.
-
OmniLiDAR: A Unified Diffusion Framework for Multi-Domain 3D LiDAR Generation
A unified text-conditioned diffusion model generates high-fidelity LiDAR scans across eight domains spanning weather, sensor, and platform shifts using cross-domain training and feature modeling.
-
Contrastive-SDXL: Annotation-Preserving Night-Time Augmentation for Pedestrian Detection
Contrastive-SDXL augments daytime images into realistic night-time versions using SDXL-Turbo with LoRA and multi-level DINOv2 contrastive losses, yielding 6-7% lower miss rate on pedestrian detection versus daytime-on...
-
GeoQuery: Geometry-Query Diffusion for Sparse-View Reconstruction
GeoQuery replaces corrupted rendering features with geometry-aligned proxy queries and restricts cross-view attention to local windows, enabling robust diffusion-based refinement under extreme view sparsity.
-
DiffNR: Diffusion-Enhanced Neural Representation Optimization for Sparse-View 3D Tomographic Reconstruction
DiffNR integrates a conditioned single-step diffusion model to generate periodic pseudo-reference volumes that provide auxiliary supervision during neural representation optimization for sparse-view tomographic recons...
-
DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
DF3DV-1K supplies 1,048 real scenes with clean/cluttered image pairs and a 41-scene hard subset to benchmark and improve distractor-free radiance-field methods.
-
RIRF: Reasoning Image Restoration Framework
R&R couples structured diagnostic reasoning from a fine-tuned Qwen3-VL model with reinforcement learning guided by degradation severity to achieve state-of-the-art universal image restoration with added interpretability.
-
GMODiff: One-Step Gain Map Refinement with Diffusion Priors for HDR Reconstruction
HDR reconstruction is reformulated as one-step gain map refinement, enabling a pre-trained latent diffusion model plus regression priors to produce high-quality HDR at a fraction of the cost of prior diffusion approaches.
-
Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding
Elastic3D converts monocular video to stereo by directly synthesizing the right-eye view with a one-step latent diffusion model conditioned on a user-set median disparity, using a guided decoder to preserve left-view details.
-
R3D2: Realistic 3D Asset Insertion via Diffusion for Autonomous Driving Simulation
R3D2 trains a lightweight diffusion model on synthetic placements of 3DGS-generated assets to produce photorealistic insertions with consistent illumination into autonomous driving scenes.
-
BeyondFusion: Self-Aligned Latent Diffusion for Calibration-Free Infrared Super-Resolution and Infrared-Visible Fusion
One latent diffusion model, with token-level cross-modal attention, performs calibration-free visible-guided infrared super-resolution and infrared-visible fusion as two outputs of the same process.
-
Cyclone: Diffusion Model for Cycle-Consistent Weather Editing from Unpaired Driving Data
A single latent-diffusion model, trained with cycle consistency, self-distillation, and CLIP guidance on unpaired driving data, edits fog, rain, and snow in driving scenes and modestly improves downstream perception.
-
FieryGS: In-the-Wild Fire Synthesis with Physics-Integrated Gaussian Splatting
FieryGS couples MLLM-based material reasoning with simplified combustion simulation and unified fire/smoke/3DGS rendering to synthesize controllable, scene-consistent fire in reconstructed real-world scenes.
-
Unposed-to-3D: Learning Simulation-Ready Vehicles from Real-World Images
Unposed-to-3D learns simulation-ready 3D vehicle models from unposed real images by predicting camera parameters for photometric self-supervision, then adding scale prediction and harmonization.
-
3D Smoke Scene Reconstruction Guided by Vision Priors from Multimodal Large Language Models
A framework that combines MLLM-based image enhancement with a medium-aware 3D Gaussian Splatting model to reconstruct and render smoke scenes.
-
Translationese as a Rational Response to Translation Task Difficulty
Translationese is partly predictable from quantifiable translation-task difficulty, especially cross-lingual transfer load, more so for English-to-German than the reverse.
-
MedShift: Implicit Conditional Transport for X-Ray Domain Adaptation
MedShift applies flow matching and Schrödinger bridges for class-conditional unpaired translation between synthetic and real skull X-rays, benchmarked on the new X-DigiSkull dataset.
-
Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution
DDR-SR routes each real-world low-resolution image to one of two diffusion experts based on a high-frequency-loss difficulty score, using a low-compression VAE for hard images and a high-compression VAE for easy image...
-
Two-Stage Cross-Domain Cervical Abnormality Screening with Cytopathological Image Synthesis and Knowledge Distillation
A two-stage framework for cross-domain cervical abnormality detection that uses Spatially-Continuous Unpaired Neural Schrödinger Bridge for image synthesis and dual-level knowledge distillation for feature alignment.
-
MariData: One-Step Unpaired Image Translation for Maritime Environments
CycleGAN-turbo with zero-convolution skip connections for unpaired maritime image translation to synthesize adverse-weather and low-light scenes while retaining small object details.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.