Pith. sign in

REVIEW 13 cited by

Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.08261 v4 pith:GS2IW2E6 submitted 2024-10-10 cs.CV

classification cs.CV
keywords meissonictext-to-imagehigh-qualityhigh-resolutionimageimageslikemasked
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

We present Meissonic, which elevates non-autoregressive masked image modeling (MIM) text-to-image to a level comparable with state-of-the-art diffusion models like SDXL. By incorporating a comprehensive suite of architectural innovations, advanced positional encoding strategies, and optimized sampling conditions, Meissonic substantially improves MIM's performance and efficiency. Additionally, we leverage high-quality training data, integrate micro-conditions informed by human preference scores, and employ feature compression layers to further enhance image fidelity and resolution. Our model not only matches but often exceeds the performance of existing models like SDXL in generating high-quality, high-resolution images. Extensive experiments validate Meissonic's capabilities, demonstrating its potential as a new standard in text-to-image synthesis. We release a model checkpoint capable of producing $1024 \times 1024$ resolution images.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Masked discrete diffusion with token editing and grouped cross-entropy reaches strong text-to-image generation scores in an 8B decoder-only model, reporting GenEval 0.90, DPG 86.9, HPSv3 10.76.

  2. IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction

    cs.CV 2025-10 conditional novelty 6.0 of 10

    IAR2 achieves state-of-the-art ImageNet 256×256 image generation (FID 1.50 with rejection sampling) by splitting visual tokens into semantic and detail codes and predicting them hierarchically with a local-context-awa...

  3. Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Lavida-O introduces an elastic mixture-of-transformers architecture that brings high-resolution text-to-image generation, object grounding, and image editing into a single masked diffusion model, using planning and se...

  4. Seeing World Dynamics in a Nutshell

    cs.CV 2025-02 conditional novelty 6.0 of 10

    NutWorld is a feed-forward model that represents a monocular video as structured dynamic 3D Gaussians in a canonical orthographic space, trained with depth and flow priors.

  5. Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A text-aware 1D tokenizer and a masked generative model, trained entirely on public data, reach FID and GenEval scores comparable to larger private-data diffusion and autoregressive models.

  6. Improving Autoregressive Visual Generation with Cluster-Oriented Token Prediction

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Rearranging a visual codebook into balanced clusters of similar codes and adding a cluster-oriented cross-entropy loss improves autoregressive image quality and training efficiency.

  7. DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DiffSensei combines an SDXL diffusion generator with a multimodal LLM adapter and masked attention to generate manga pages with multiple characters whose poses and expressions follow panel captions.

  8. HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing

    cs.CV 2024-12 conditional novelty 6.0 of 10

    HumanEdit provides 5,751 human-annotated, high-resolution image editing pairs with masks and a six-type instruction taxonomy, plus baseline benchmark results.

  9. The Efficacy of Transfer-based No-box Attacks on Image Watermarking: A Pragmatic Analysis

    cs.CR 2024-12 conditional novelty 6.0 of 10

    Transfer-based no-box watermark evasion largely fails without aligned surrogate models, and a simple one-surrogate perturbation (OFT) matches or exceeds the expensive optimization-based attack in 11 of 12 tested confi...

  10. AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A large automatically collected image editing dataset with 25 editing types and a task-aware diffusion model trained on it achieve new state-of-the-art results on two standard image editing benchmarks.

  11. Bag of Design Choices for Inference of High-Resolution Masked Generative Transformer

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A set of inference-time design choices improves image quality, memory, and speed of masked generative Transformers, with combined tricks winning about 70% of human-preference comparisons against vanilla sampling.

  12. DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer

    cs.CV 2025-07 conditional novelty 5.0 of 10

    DC-AR generates 512x512 images in 12 masked autoregressive steps plus 20 diffusion refinement steps, using a 32x compressed 2D tokenizer, and reports gFID 5.49 on MJHQ-30K.

  13. Track Any Anomalous Object: A Granular Video Anomaly Detection Pipeline

    cs.CV 2025-06 conditional novelty 5.0 of 10

    TAO pipelines object-centric anomaly scores into SAM2 prompts with a temporal consistency filter to obtain pixel-level anomaly segmentation and tracking.

Pith tools