Pith. sign in

REVIEW 3 cited by

AccDiffusion: An Accurate Method for Higher-Resolution Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.10738 v2 pith:J2UKIXQQ submitted 2024-07-15 cs.CV

classification cs.CV
keywords generationimageaccdiffusionhigher-resolutionobjectpromptaccuratebetter
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper attempts to address the object repetition issue in patch-wise higher-resolution image generation. We propose AccDiffusion, an accurate method for patch-wise higher-resolution image generation without training. An in-depth analysis in this paper reveals an identical text prompt for different patches causes repeated object generation, while no prompt compromises the image details. Therefore, our AccDiffusion, for the first time, proposes to decouple the vanilla image-content-aware prompt into a set of patch-content-aware prompts, each of which serves as a more precise description of an image patch. Besides, AccDiffusion also introduces dilated sampling with window interaction for better global consistency in higher-resolution image generation. Experimental comparison with existing methods demonstrates that our AccDiffusion effectively addresses the issue of repeated object generation and leads to better performance in higher-resolution image generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CineScale: Free Lunch in High-Resolution Cinematic Visual Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    CineScale extends pre-trained diffusion models to 8k image and 4k video generation with mostly tuning-free inference plus a small LoRA adaptation for video.

  2. FreeScale: Unleashing the Resolution of Diffusion Models via Tuning-Free Scale Fusion

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A tuning-free scale-fusion method that lets frozen diffusion models generate 8k images and high-res videos by combining global and local attention through frequency filtering.

  3. Parallel Sequence Modeling via Generalized Spatial Propagation Network

    cs.CV 2025-01 conditional novelty 5.0 of 10

    GSPN is a 2D line-scan propagation mechanism for vision that reports SOTA ImageNet accuracy, strong class-conditional generation FID, and large high-resolution text-to-image speedups.

Pith tools