REVIEW 13 cited by
Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
We present Meissonic, which elevates non-autoregressive masked image modeling (MIM) text-to-image to a level comparable with state-of-the-art diffusion models like SDXL. By incorporating a comprehensive suite of architectural innovations, advanced positional encoding strategies, and optimized sampling conditions, Meissonic substantially improves MIM's performance and efficiency. Additionally, we leverage high-quality training data, integrate micro-conditions informed by human preference scores, and employ feature compression layers to further enhance image fidelity and resolution. Our model not only matches but often exceeds the performance of existing models like SDXL in generating high-quality, high-resolution images. Extensive experiments validate Meissonic's capabilities, demonstrating its potential as a new standard in text-to-image synthesis. We release a model checkpoint capable of producing $1024 \times 1024$ resolution images.
Forward citations
Cited by 13 Pith papers
-
Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis
Masked discrete diffusion with token editing and grouped cross-entropy reaches strong text-to-image generation scores in an 8B decoder-only model, reporting GenEval 0.90, DPG 86.9, HPSv3 10.76.
-
IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction
IAR2 achieves state-of-the-art ImageNet 256×256 image generation (FID 1.50 with rejection sampling) by splitting visual tokens into semantic and detail codes and predicting them hierarchically with a local-context-awa...
-
Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
Lavida-O introduces an elastic mixture-of-transformers architecture that brings high-resolution text-to-image generation, object grounding, and image editing into a single masked diffusion model, using planning and se...
-
Seeing World Dynamics in a Nutshell
NutWorld is a feed-forward model that represents a monocular video as structured dynamic 3D Gaussians in a canonical orthographic space, trained with depth and flow priors.
-
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
A text-aware 1D tokenizer and a masked generative model, trained entirely on public data, reach FID and GenEval scores comparable to larger private-data diffusion and autoregressive models.
-
Improving Autoregressive Visual Generation with Cluster-Oriented Token Prediction
Rearranging a visual codebook into balanced clusters of similar codes and adding a cluster-oriented cross-entropy loss improves autoregressive image quality and training efficiency.
-
DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation
DiffSensei combines an SDXL diffusion generator with a multimodal LLM adapter and masked attention to generate manga pages with multiple characters whose poses and expressions follow panel captions.
-
HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing
HumanEdit provides 5,751 human-annotated, high-resolution image editing pairs with masks and a six-type instruction taxonomy, plus baseline benchmark results.
-
The Efficacy of Transfer-based No-box Attacks on Image Watermarking: A Pragmatic Analysis
Transfer-based no-box watermark evasion largely fails without aligned surrogate models, and a simple one-surrogate perturbation (OFT) matches or exceeds the expensive optimization-based attack in 11 of 12 tested confi...
-
AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
A large automatically collected image editing dataset with 25 editing types and a task-aware diffusion model trained on it achieve new state-of-the-art results on two standard image editing benchmarks.
-
Bag of Design Choices for Inference of High-Resolution Masked Generative Transformer
A set of inference-time design choices improves image quality, memory, and speed of masked generative Transformers, with combined tricks winning about 70% of human-preference comparisons against vanilla sampling.
-
DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer
DC-AR generates 512x512 images in 12 masked autoregressive steps plus 20 diffusion refinement steps, using a 32x compressed 2D tokenizer, and reports gFID 5.49 on MJHQ-30K.
-
Track Any Anomalous Object: A Granular Video Anomaly Detection Pipeline
TAO pipelines object-centric anomaly scores into SAM2 prompts with a temporal consistency filter to obtain pixel-level anomaly segmentation and tracking.
Discussion (0). Continue with ORCID to comment.