REVIEW 3 cited by
Diffusion-based Aesthetic QR Code Generation via Scanning-Robust Perceptual Guidance
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Diffusion-based Aesthetic QR Code Generation via Scanning-Robust Perceptual Guidance
read the original abstract
QR codes, prevalent in daily applications, lack visual appeal due to their conventional black-and-white design. Integrating aesthetics while maintaining scannability poses a challenge. In this paper, we introduce a novel diffusion-model-based aesthetic QR code generation pipeline, utilizing pre-trained ControlNet and guided iterative refinement via a novel classifier guidance (SRG) based on the proposed Scanning-Robust Loss (SRL) tailored with QR code mechanisms, which ensures both aesthetics and scannability. To further improve the scannability while preserving aesthetics, we propose a two-stage pipeline with Scanning-Robust Perceptual Guidance (SRPG). Moreover, we can further enhance the scannability of the generated QR code by post-processing it through the proposed Scanning-Robust Projected Gradient Descent (SRPGD) post-processing technique based on SRL with proven convergence. With extensive quantitative, qualitative, and subjective experiments, the results demonstrate that the proposed approach can generate diverse aesthetic QR codes with flexibility in detail. In addition, our pipelines outperforming existing models in terms of Scanning Success Rate (SSR) 86.67% (+40%) with comparable aesthetic scores. The pipeline combined with SRPGD further achieves 96.67% (+50%). Our code will be available https://github.com/jwliao1209/DiffQRCode.
Forward citations
Cited by 3 Pith papers
-
Text-to-Image Generation for Projector-Camera System Registration
Text-conditioned images with controlled ORB-feature distributions plus a distortion-trained registration network enable accurate, natural-image procam registration on smooth surfaces.
-
Structure-Semantic Co-optimized Latent Diffusion Model for Fast Visual Anagram Synthesis
S2CO-Anagram adapts multi-view parallel denoising to SDXL-Turbo with null-text structure alignment, semantic enhancement, and attention-guided noise fusion to produce superior 512x512 visual anagrams in ~2.6s.
-
Structure-Semantic Co-optimized Latent Diffusion Model for Fast Visual Anagram Synthesis
A structure-semantic co-optimization framework on few-step latent diffusion produces higher-resolution visual anagrams with better fidelity and much lower inference time than prior SOTA.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.