Pith. sign in

REVIEW 2 cited by

Unified Discrete Diffusion for Simultaneous Vision-Language Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.14842 v1 pith:GO4DU2VN submitted 2022-11-27 cs.CV

classification cs.CV
keywords generationunifieddiffusiondiscretemulti-modalitymodelmultimodalperform
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recently developed discrete diffusion models perform extraordinarily well in the text-to-image task, showing significant promise for handling the multi-modality signals. In this work, we harness these traits and present a unified multimodal generation model that can conduct both the "modality translation" and "multi-modality generation" tasks using a single model, performing text-based, image-based, and even vision-language simultaneous generation. Specifically, we unify the discrete diffusion process for multimodal signals by proposing a unified transition matrix. Moreover, we design a mutual attention module with fused embedding layer and a unified objective function to emphasise the inter-modal linkages, which are vital for multi-modality generation. Extensive experiments indicate that our proposed method can perform comparably to the state-of-the-art solutions in various generation tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A unified diffusion framework with per-modality noise clocks lets one model generate images, text, and tabular data jointly or conditionally in their native spaces.

  2. FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A 1.5B unified multimodal model trained with discrete flow matching and metric-induced probability paths matches autoregressive baselines of similar size on generation and understanding benchmarks.

Pith tools