REVIEW 19 cited by
Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Image-to-image translation is a class of vision and graphics problems where the goal is to learn the mapping between an input image and an output image using a training set of aligned image pairs. However, for many tasks, paired training data will not be available. We present an approach for learning to translate an image from a source domain $X$ to a target domain $Y$ in the absence of paired examples. Our goal is to learn a mapping $G: X \rightarrow Y$ such that the distribution of images from $G(X)$ is indistinguishable from the distribution $Y$ using an adversarial loss. Because this mapping is highly under-constrained, we couple it with an inverse mapping $F: Y \rightarrow X$ and introduce a cycle consistency loss to push $F(G(X)) \approx X$ (and vice versa). Qualitative results are presented on several tasks where paired training data does not exist, including collection style transfer, object transfiguration, season transfer, photo enhancement, etc. Quantitative comparisons against several prior methods demonstrate the superiority of our approach.
Forward citations
Cited by 19 Pith papers
-
FLAME 3 Dataset: Unleashing the Power of Radiometric Thermal UAV Imagery for Wildfire Management
FLAME 3 provides the first aerial radiometric thermal wildfire image dataset with per-pixel temperature TIFFs and nadir thermal plots, plus a pipeline and benchmark.
-
Bridging Scales in Map Generation: A scale-aware cascaded generative mapping framework for seamless and consistent multi-scale cartographic representation
A cascaded latent-diffusion framework with CLIP-based scale encoding and cascade references generates seamless multi-scale tile maps from remote sensing imagery, reporting state-of-the-art FID/PSNR on MLMG and CSCMG b...
-
Domain Intersection and Domain Difference
A symmetric encoder-decoder with zero, adversarial, and reconstruction losses disentangles shared and domain-specific content, enabling guided image translation and generation of intersection and union domains without...
-
Improved MR to CT synthesis for PET/MR attenuation correction using Imitation Learning
A multi-hypothesis CNN with a learned PET-residual loss produces pseudo-CTs that reduce PET reconstruction error at the cost of higher CT error.
-
Deep Tone Mapping Operator for High Dynamic Range Images
DeepTMO is a multi-scale conditional GAN that tone-maps 32-bit HDR images to high-resolution LDR outputs, reporting higher TMQI scores and a subjective preference over classical TMOs.
-
Adversarial Self-Defense for Cycle-Consistent GANs
Two defenses for cycle-consistent GANs, additive noise and a guess discriminator, reduce hidden-information 'self-adversarial attacks' and improve robustness to high-frequency perturbations.
-
Transformer-based Diffusion models for Hydrological Time Series Probabilistic Imputation and Forecasting
A modified CSDI diffusion model with convolutions, RMSNorm, and Fourier encoding jointly imputes and forecasts hydrological time series, outperforming standard baselines on two datasets for short horizons.
-
Counterfactual Explainability Framework With CycleGAN And Counterfactual-Classifier Alignnment Score for Retinal Disease Classification
A CycleGAN-based counterfactual framework translates diseased retinal images to healthy-looking counterparts, and a new CCAS metric scores spatial agreement between the translation difference maps and classifier saliency.
-
Underwater Monocular Metric Depth Estimation: Real-World Benchmarks and Synthetic Fine-Tuning with Vision Foundation Models
Fine-tuning Depth Anything V2 on physics-based synthetic underwater versions of Hypersim improves metric depth accuracy on real underwater benchmarks like FLSea and SQUID, though one AbsRel number worsens slightly.
-
Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis
A video diffusion model, HunyuanVideo-I2V, is adapted with mixup transitions, frame-skip position embeddings, and attention masking to outperform image-only models on several controllable image generation benchmarks.
-
Learning more with the same effort: how randomization improves the robustness of a robotic deep reinforcement learning agent
Randomizing camera position during simulated robot-arm training improves robustness to viewpoint changes by about 25 percent average accuracy over fixed-camera training, at the same training budget.
-
PyPotteryLens: An Open-Source Deep Learning Framework for Automated Digitisation of Archaeological Pottery Documentation
PyPotteryLens detects, segments, orients, and labels pottery drawings from archaeological PDFs using YOLO and EfficientNetV2, reporting above 96% precision and up to 20x faster processing.
-
A Tour of Convolutional Networks Guided by Linear Interpreters
A hooking layer (LinearScope) freezes the nonlinear decisions of a CNN to expose the network as a single linear map, revealing bias-dominated classifier scores, wavelet-like super-resolution bases, and copy-move/templ...
-
ADN: Artifact Disentanglement Network for Unsupervised Metal Artifact Reduction
ADN, an artifact disentanglement network, reduces CT metal artifacts using only unpaired artifact-free and artifact-affected images, matching supervised methods on synthesized data.
-
Unsupervised Domain Adaptation for Multitask Image Analysis in Realistic Context with Extreme Label Shift; Application to the CTAO first Large Sized Telescope
A comparative study showing that conditional domain adaptation combined with uncertainty-based multitask balancing improves gamma-ray event reconstruction under extreme label shift in CTAO LST simulations.
-
SSDD-GAN: Single-Step Denoising Diffusion GAN for Cochlear Implant Surgical Scene Completion
A single-step denoising diffusion GAN with a Patch-GAN discriminator completes surgical microscope scenes, reporting higher SSIM than several inpainting baselines on a small single-patient dataset.
-
Learning Deep Representations by Mutual Information for Person Re-identification
Adding a Deep InfoMax-style adversarial loss to IDE and PCB person re-identification baselines gives modest rank-1/mAP gains, but the claimed mutual information between input image and encoder output is not implemente...
-
What goes around comes around: Cycle-Consistency-based Short-Term Motion Prediction for Anomaly Detection using Generative Adversarial Networks
A cross-channel GAN for video anomaly detection improves UCSD Ped2 AUC from 93.7% to 98.0% by adding a cycle-consistency loss and morphological noise suppression.
-
Learning Text Styles: A Study on Transfer, Attribution, and Verification
A thesis compiles published work claiming that lightweight adapters, contrastive disentanglement, and instruction tuning improve text style transfer, authorship attribution, and authorship verification.
Discussion (0). Continue with ORCID to comment.