Pith. sign in

REVIEW 7 cited by

Text-to-Image Rectified Flow as Plug-and-Play Priors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.03293 v4 pith:7WMH5FJK submitted 2024-06-05 cs.CV

Text-to-Image Rectified Flow as Plug-and-Play Priors

classification cs.CV
keywords modelsrectifieddiffusionflowpriorsgenerativeperformancegeneration
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Large-scale diffusion models have achieved remarkable performance in generative tasks. Beyond their initial training applications, these models have proven their ability to function as versatile plug-and-play priors. For instance, 2D diffusion models can serve as loss functions to optimize 3D implicit models. Rectified flow, a novel class of generative models, enforces a linear progression from the source to the target distribution and has demonstrated superior performance across various domains. Compared to diffusion-based methods, rectified flow approaches surpass in terms of generation quality and efficiency, requiring fewer inference steps. In this work, we present theoretical and experimental evidence demonstrating that rectified flow based methods offer similar functionalities to diffusion models - they can also serve as effective priors. Besides the generative capabilities of diffusion priors, motivated by the unique time-symmetry properties of rectified flow models, a variant of our method can additionally perform image inversion. Experimentally, our rectified flow-based priors outperform their diffusion counterparts - the SDS and VSD losses - in text-to-3D generation. Our method also displays competitive performance in image inversion and editing.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. $Z^2$-Sampling: Zero-Cost Zigzag Trajectories for Semantic Alignment in Diffusion Models

    cs.CV 2026-04 unverdicted novelty 7.0

    Z²-Sampling implicitly realizes zero-cost zigzag trajectories for curvature-aware semantic alignment in diffusion models by reducing multi-step paths via operator dualities and temporal caching while synthesizing a di...

  2. Flow Straight and Fast in Hilbert Space: Functional Rectified Flow

    cs.LG 2025-09 conditional novelty 7.0

    Functional rectified flow is defined and proved to preserve marginals in separable Hilbert spaces, with functional flow matching and probability-flow ODEs as special cases.

  3. JanusMesh: Fast and Zero-Shot 3D Visual Illusion Generation via Cross-Space Denoising

    cs.CV 2026-06 unverdicted novelty 6.0

    A training-free two-stage pipeline uses cross-space dual-branch denoising with CLIP-guided voxel alignment and SDF blending for geometry, followed by view-conditioned 2D diffusion texture projection, to produce dual-s...

  4. VAGS: Velocity Adaptive Guidance Scale for Image Editing and Generation

    cs.CV 2026-05 accept novelty 6.0

    VAGS adapts the CFG scale at each ODE step using velocity alignment signals to raise structural fidelity in editing and sample quality in generation over fixed-scale baselines.

  5. Translationese as a Rational Response to Translation Task Difficulty

    cs.CL 2026-03 unverdicted novelty 5.0

    Translationese is partly predictable from quantifiable translation-task difficulty, especially cross-lingual transfer load, more so for English-to-German than the reverse.

  6. FlowSteer: Conditioning Flow Field for Consistent Image Restoration

    eess.IV 2025-12 conditional novelty 5.0

    A sparse mid-to-late schedule of null-space fidelity updates lets a frozen text-to-image flow model restore images with high measurement consistency.

  7. FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration

    cs.SD 2026-07 reject novelty 4.0

    FlowSonic combines deterministic rectified-flow inversion, cached cross-attention injection, and a 'seeded' third-order Adams-Bashforth solver to report better timbre and genre edits on small datasets.