Pith. sign in

REVIEW 2 cited by

Multi-Scale Diffusion: Enhancing Spatial Layout in High-Resolution Panoramic Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.18830 v2 pith:DZFIB4EY submitted 2024-10-24 cs.CV

classification cs.CV
keywords high-resolutionimagediffusiongenerationimageslayoutpanoramicframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models have recently gained recognition for generating diverse and high-quality content, especially in image synthesis. These models excel not only in creating fixed-size images but also in producing panoramic images. However, existing methods often struggle with spatial layout consistency when producing high-resolution panoramas due to the lack of guidance on the global image layout. This paper introduces the Multi-Scale Diffusion (MSD), an optimized framework that extends the panoramic image generation framework to multiple resolution levels. Our method leverages gradient descent techniques to incorporate structural information from low-resolution images into high-resolution outputs. Through comprehensive qualitative and quantitative evaluations against prior work, we demonstrate that our approach significantly improves the coherence of high-resolution panorama generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A shared-backbone multi-scale diffusion transformer with motion-score conditioning generates competitive videos at reduced compute and with adjustable dynamics.

  2. PanoLlama: Generating Endless and Coherent Panoramas with Next-Token-Prediction LLMs

    cs.CV 2024-11 conditional novelty 6.0 of 10

    PanoLlama uses token redirection on a fixed-size autoregressive image model to generate coherent, arbitrarily long panoramas without any extra training.

Pith tools