Pith. sign in

REVIEW 4 cited by

OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.01169 v2 pith:IUY7P6GJ submitted 2024-12-02 cs.MM cs.CVcs.SDeess.AS

classification cs.MMcs.CVcs.SDeess.AS
keywords text-to-imagegenerationany-to-anymodalitiesnovelomniflowrectifiedarchitecture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce OmniFlow, a novel generative model designed for any-to-any generation tasks such as text-to-image, text-to-audio, and audio-to-image synthesis. OmniFlow advances the rectified flow (RF) framework used in text-to-image models to handle the joint distribution of multiple modalities. It outperforms previous any-to-any models on a wide range of tasks, such as text-to-image and text-to-audio synthesis. Our work offers three key contributions: First, we extend RF to a multi-modal setting and introduce a novel guidance mechanism, enabling users to flexibly control the alignment between different modalities in the generated outputs. Second, we propose a novel architecture that extends the text-to-image MMDiT architecture of Stable Diffusion 3 and enables audio and text generation. The extended modules can be efficiently pretrained individually and merged with the vanilla text-to-image MMDiT for fine-tuning. Lastly, we conduct a comprehensive study on the design choices of rectified flow transformers for large-scale audio and text generation, providing valuable insights into optimizing performance across diverse modalities. The Code will be available at https://github.com/jacklishufan/OmniFlows.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Flow Straight and Fast in Hilbert Space: Functional Rectified Flow

    cs.LG 2025-09 conditional novelty 7.0 of 10

    Functional rectified flow is defined and proved to preserve marginals in separable Hilbert spaces, with functional flow matching and probability-flow ODEs as special cases.

  2. Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Lavida-O introduces an elastic mixture-of-transformers architecture that brings high-resolution text-to-image generation, object grounding, and image editing into a single masked diffusion model, using planning and se...

  3. Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A unified diffusion framework with per-modality noise clocks lets one model generate images, text, and tabular data jointly or conditionally in their native spaces.

  4. Flow Diverse and Efficient: Learning Momentum Flow Matching via Stochastic Velocity Field Sampling

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Momentum Flow perturbs rectified flow velocities with a decaying random component and shows improved FID and recall on CelebA-HQ with half the sampling steps.

Pith tools