Pith. sign in

REVIEW 2 cited by

FlowDCN: Exploring DCN-like Architectures for Fast Image Generation with Arbitrary Resolution

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.22655 v1 pith:4GVFKEMD submitted 2024-10-30 cs.CV

classification cs.CV
keywords flowdcnimageresolutionresolutionsarbitraryextrapolationgenerationimages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Arbitrary-resolution image generation still remains a challenging task in AIGC, as it requires handling varying resolutions and aspect ratios while maintaining high visual quality. Existing transformer-based diffusion methods suffer from quadratic computation cost and limited resolution extrapolation capabilities, making them less effective for this task. In this paper, we propose FlowDCN, a purely convolution-based generative model with linear time and memory complexity, that can efficiently generate high-quality images at arbitrary resolutions. Equipped with a new design of learnable group-wise deformable convolution block, our FlowDCN yields higher flexibility and capability to handle different resolutions with a single model. FlowDCN achieves the state-of-the-art 4.30 sFID on $256\times256$ ImageNet Benchmark and comparable resolution extrapolation results, surpassing transformer-based counterparts in terms of convergence speed (only $\frac{1}{5}$ images), visual quality, parameters ($8\%$ reduction) and FLOPs ($20\%$ reduction). We believe FlowDCN offers a promising solution to scalable and flexible image synthesis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Native-Resolution Image Synthesis

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A single diffusion transformer trained on native-resolution ImageNet achieves state-of-the-art FID at 256 and 512, and extrapolates to 1024 and 1536 with moderate degradation.

  2. Differentiable Solver Search for Fast Diffusion Sampling

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A differentiable search over solver coefficients and sampling timesteps produces a fast diffusion sampler that outperforms DPM-Solver++ and UniPC at 5 to 10 steps.

Pith tools