Pith. sign in

REVIEW 5 cited by

aMUSEd: An Open MUSE Reproduction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.01808 v1 pith:AJWBQZFE submitted 2024-01-03 cs.CV

classification cs.CV
keywords generationamusedimagemusetext-to-imagecompareddiffusionlatent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present aMUSEd, an open-source, lightweight masked image model (MIM) for text-to-image generation based on MUSE. With 10 percent of MUSE's parameters, aMUSEd is focused on fast image generation. We believe MIM is under-explored compared to latent diffusion, the prevailing approach for text-to-image generation. Compared to latent diffusion, MIM requires fewer inference steps and is more interpretable. Additionally, MIM can be fine-tuned to learn additional styles with only a single image. We hope to encourage further exploration of MIM by demonstrating its effectiveness on large-scale text-to-image generation and releasing reproducible training code. We also release checkpoints for two models which directly produce images at 256x256 and 512x512 resolutions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models

    cs.CV 2025-12 conditional novelty 7.0 of 10

    dMLLM-TTS delivers up to 6x more efficient test-time scaling for diffusion MLLMs via O(N+T) hierarchical search and self-verified feedback, improving generation quality on GenEval across three models.

  2. Controllable Image Generation with Composed Parallel Token Prediction

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    A new formulation for composing discrete generative processes enables precise control over novel condition combinations in image generation, cutting error rates by 63% and speeding up inference.

  3. Controllable Image Generation with Composed Parallel Token Prediction

    cs.CV 2024-05 unverdicted novelty 6.0 of 10

    A derived formulation for composing discrete probabilistic generative processes enables novel condition combinations in image generation, yielding 63.4% relative error reduction and FID gains on CLEVR and FFHQ datasets.

  4. Dissect and Prune: Enhancing Robustness in AI-Generated Image Detection

    cs.CV 2026-06 unverdicted novelty 5.0 of 10

    DEAR prunes channel features whose activations align strongly with inpaint masks, retaining only those capturing genuine generative artifacts to improve robustness against post-processing and unseen generators.

  5. World Model on Million-Length Video And Language With Blockwise RingAttention

    cs.LG 2024-02 unverdicted novelty 5.0 of 10

    Presents open-source 7B models for million-token video and language understanding via Blockwise RingAttention, setting new benchmarks in retrieval and long video tasks.

Pith tools