Pith. sign in

REVIEW 38 cited by

Taming Rectified Flow for Inversion and Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.04746 v3 pith:C2YITPAK submitted 2024-11-07 cs.CV

Taming Rectified Flow for Inversion and Editing

classification cs.CV
keywords editingimagevideoinversionflowprocessrectifiedtasks
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Rectified-flow-based diffusion transformers like FLUX and OpenSora have demonstrated outstanding performance in the field of image and video generation. Despite their robust generative capabilities, these models often struggle with inversion inaccuracies, which could further limit their effectiveness in downstream tasks such as image and video editing. To address this issue, we propose RF-Solver, a novel training-free sampler that effectively enhances inversion precision by mitigating the errors in the ODE-solving process of rectified flow. Specifically, we derive the exact formulation of the rectified flow ODE and apply the high-order Taylor expansion to estimate its nonlinear components, significantly enhancing the precision of ODE solutions at each timestep. Building upon RF-Solver, we further propose RF-Edit, a general feature-sharing-based framework for image and video editing. By incorporating self-attention features from the inversion process into the editing process, RF-Edit effectively preserves the structural information of the source image or video while achieving high-quality editing results. Our approach is compatible with any pre-trained rectified-flow-based models for image and video tasks, requiring no additional training or optimization. Extensive experiments across generation, inversion, and editing tasks in both image and video modalities demonstrate the superiority and versatility of our method. The source code is available at https://github.com/wangjiangshan0725/RF-Solver-Edit.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 38 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Improving Robotic Generalist Policies via Flow Reversal Steering

    cs.RO 2026-06 unverdicted novelty 7.0

    Flow Reversal Steering steers flow matching generalist policies by reversing suboptimal actions to nearby better modes, enabling improved zero-shot control, quick distillation, and RL bootstrapping in robotic manipulation.

  2. Mixture Prototype Flow Matching for Open-Set Supervised Anomaly Detection

    cs.CV 2026-05 unverdicted novelty 7.0

    MPFM uses flow matching with a Gaussian mixture prior on the velocity field and a mutual information maximizer to improve open-set anomaly detection over unimodal prototype methods.

  3. Mixture Prototype Flow Matching for Open-Set Supervised Anomaly Detection

    cs.CV 2026-05 unverdicted novelty 7.0

    MPFM models flow matching velocity as a Gaussian mixture prior per normal class plus a mutual information regularizer to improve open-set anomaly detection over unimodal prototypes.

  4. DirectEdit: Step-Level Accurate Inversion for Flow-Based Image Editing

    cs.CV 2026-05 unverdicted novelty 7.0

    DirectEdit achieves step-level accurate inversion for flow-based image editing by directly aligning forward paths, using attention feature injection and mask-guided noise blending to balance fidelity and editability w...

  5. UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models

    cs.CV 2026-04 unverdicted novelty 7.0

    UniGeo unifies geometric guidance across three levels in video models to reduce geometric drift and improve consistency in camera-controllable image editing.

  6. UniEditBench: A Unified and Cost-Effective Benchmark for Image and Video Editing via Distilled MLLMs

    cs.CV 2026-04 unverdicted novelty 7.0

    UniEditBench unifies image and video editing evaluation with a nine-plus-eight operation taxonomy and cost-effective 4B/8B distilled MLLM evaluators that align with human judgments.

  7. Prompt-Guided Image Editing with Masked Logit Nudging in Visual Autoregressive Models

    cs.CV 2026-04 unverdicted novelty 7.0

    Masked Logit Nudging aligns visual autoregressive model logits with source token maps under target prompts inside cross-attention masks, delivering top image editing results on PIE benchmarks and strong reconstruction...

  8. VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space

    cs.CV 2025-08 conditional novelty 7.0

    A training-free 3D editing method that inverts a source asset into TRELLIS latent space and replaces latents plus attention K/V tokens in unedited regions during re-denosing.

  9. In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer

    cs.CV 2025-04 unverdicted novelty 7.0

    ICEdit achieves state-of-the-art instructional image editing in Diffusion Transformers via in-context generation, requiring only 0.1% of prior training data and 1% trainable parameters.

  10. UniEdit-Flow: Unleashing Inversion and Editing in the Era of Flow Models

    cs.CV 2025-04 unverdicted novelty 7.0

    UniEdit-Flow presents tuning-free Uni-Inv and Uni-Edit methods for inversion and editing in flow models that achieve accurate reconstruction and robust region-preserving edits across generative models.

  11. LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing

    cs.CV 2026-06 conditional novelty 6.5

    A three-stage distillation plus AR mask cache converts a bidirectional DiT editor into a real-time causal streaming editor that preserves non-edited regions at 12.66 FPS.

  12. LiveLight: Real-time Streaming Video Relighting with Interactive Control

    cs.CV 2026-08 conditional novelty 6.0

    A diffusion-based system performs real-time, interactive video relighting by injecting multi-plane light irradiance conditions and streaming latent chunks.

  13. GeoEdit: Geometry-Aware Object Editing via Dual-Branch Denoising

    cs.CV 2026-06 unverdicted novelty 6.0

    GeoEdit introduces a Lift-Manipulate-Render-Denoise pipeline with dual-branch denoising and variance-homogeneous injection for 3D-consistent object editing in single photos.

  14. LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing

    cs.CV 2026-06 unverdicted novelty 6.0

    LiveEdit distills a bidirectional video foundation model into a unidirectional streaming editor via three-stage training plus mask caching to reach 12.66 FPS with stable edits.

  15. scCBGM: Interpretable Single-Cell Counterfactual Editing

    cs.LG 2026-06 unverdicted novelty 6.0

    scCBGM adapts concept bottleneck generative models with skip connections and cross-covariance penalties for single-cell data, enabling interpretable counterfactual editing and showing superior combinatorial generaliza...

  16. StreamEdit: Training-Free Video Editing via Few-Step Streaming Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    StreamEdit enables high-quality training-free video editing by adapting streaming video generation models with dual-branch fast sampling, self-attention bridge, cross-attention grounding, source-oriented guidance, and...

  17. StreamEdit: Training-Free Video Editing via Few-Step Streaming Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    StreamGVE enables high-quality training-free video editing by converting the task to noise-to-data streaming generation with dual-branch fast sampling, self-attention bridges, cross-attention grounding, source-oriente...

  18. Semantic Granularity Navigation in Image Editing

    cs.CV 2026-05 unverdicted novelty 6.0

    NaviEdit reallocates fixed step budgets in diffusion rollouts to intermediate scales for improved semantic editability while preserving fidelity via a self-consistency contract.

  19. Controlla: Learning Controllability via Graph-Constrained Latent Geometry

    cs.CV 2026-05 unverdicted novelty 6.0

    Controlla learns identity and attribute factors from multimodal inputs and aligns them with graph priors using graph-constrained optimal transport to enforce consistent attribute trajectories while preserving referenc...

  20. VAGS: Velocity Adaptive Guidance Scale for Image Editing and Generation

    cs.CV 2026-05 accept novelty 6.0

    VAGS adapts the CFG scale at each ODE step using velocity alignment signals to raise structural fidelity in editing and sample quality in generation over fixed-scale baselines.

  21. LimeCross: Context-Conditioned Layered Image Editing with Structural Consistency

    cs.CV 2026-05 unverdicted novelty 6.0

    LimeCross enables text-guided editing of individual layers in composite images by conditioning on cross-layer context via bi-stream attention while preserving layer integrity and introducing the LayerEditBench benchmark.

  22. Mixture Prototype Flow Matching for Open-Set Supervised Anomaly Detection

    cs.CV 2026-05 unverdicted novelty 6.0

    MPFM transforms normal features into a structured Gaussian mixture prototype space via a mixture velocity field and mutual information regularization to achieve state-of-the-art open-set supervised anomaly detection.

  23. FluSplat: Sparse-View 3D Editing without Test-Time Optimization

    cs.CV 2026-04 unverdicted novelty 6.0

    FluSplat trains a model with geometric alignment constraints on multi-view edits to produce consistent 3D scene edits from sparse views in a single forward pass without test-time optimization.

  24. UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models

    cs.CV 2026-04 unverdicted novelty 6.0

    UniGeo adds unified geometric guidance at three levels in video models to reduce geometric drift and improve structural fidelity in camera-controllable image editing.

  25. FlowID : Enhancing Forensic Identification with Latent Flow-Matching Models

    cs.CV 2026-03 conditional novelty 6.0

    FlowID combines single-image fine-tuning of Stable Diffusion 3 with attention-derived masks to remove death-related facial artifacts while preserving identity, outperforming open-source editors on the new InjuredFaces...

  26. RL-RIG: A Generative Spatial Reasoner via Intrinsic Reflection

    cs.CV 2026-02 unverdicted novelty 6.0

    RL-RIG uses a generate-reflect-edit loop with reinforcement learning to improve spatial accuracy in image generation, reporting up to 11% gains over prior open-source models on scene-graph metrics.

  27. CNS-Edit++: Category-Agnostic 3D Editing with Coupled Neural Shape Representation

    cs.CV 2026-07 conditional novelty 5.0

    Coupling a global latent code with a 3D feature volume lets off-the-shelf 3D generators perform local semantic edits — copy, delete, resize, mix, and drag — across object categories while preserving unedited regions.

  28. MagicPrompt: Ultra-Lightweight Prompt Tuning for Video Generation

    cs.CV 2026-07 conditional novelty 5.0

    Soft prompts inserted into attention key/value streams plus a dual pixel/latent reward let video diffusion models be tuned with under 1% trainable parameters at competitive quality.

  29. FDM-MFVT: Few-step Sampling Diffusion Model for Mask-Free Virtual Try-On

    cs.CV 2026-06 unverdicted novelty 5.0

    FDM-MFVT is a few-step mask-free virtual try-on diffusion model using OANO and IDT modules plus a new 30,000-pair MFVT dataset, claiming better efficiency and quality than baselines.

  30. EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation

    cs.CV 2026-05 unverdicted novelty 5.0

    EasyVFX decouples VFX generation via frequency-aware Mixture-of-Experts and test-time training to achieve realistic effects with limited resources.

  31. Embedding-perturbed Exploration Preference Optimization for Flow Models

    cs.CV 2026-05 unverdicted novelty 5.0

    E²PO uses embedding-level perturbations to maintain intra-group variance and discriminative signal in RL-based preference optimization for generative flow models.

  32. Semantic-Structural Alignment for Generative Pictorial Charts

    cs.GR 2026-05 unverdicted novelty 5.0

    Dual-conditioned Multi-Modal Diffusion Transformer with structural and semantic alignment mechanisms generates pictorial charts from text prompts and abstract chart images.

  33. DirectEdit: Step-Level Accurate Inversion for Flow-Based Image Editing

    cs.CV 2026-05 unverdicted novelty 5.0

    DirectEdit eliminates reconstruction error in flow-based image editing by aligning forward paths and applying attention feature injection with mask-guided noise blending.

  34. UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models

    cs.CV 2026-04 conditional novelty 5.0

    UniGeo improves camera-controllable image editing by injecting point cloud geometry into a video diffusion model at the representation, architecture, and loss levels, achieving state-of-the-art geometric consistency o...

  35. FlowSteer: Conditioning Flow Field for Consistent Image Restoration

    eess.IV 2025-12 conditional novelty 5.0

    A sparse mid-to-late schedule of null-space fidelity updates lets a frozen text-to-image flow model restore images with high measurement consistency.

  36. Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent

    cs.CV 2025-08 conditional novelty 5.0

    DescriptiveEdit turns semantic editing into reference-conditioned text-to-image generation, reporting state-of-the-art scores on the Emu Edit benchmark with a frozen backbone and about 75M trainable parameters.

  37. Semantic Granularity Navigation in Image Editing

    cs.CV 2026-05 unverdicted novelty 4.0

    NaviEdit is a training-free inference-time controller that decouples edit progress from model scale traversal in diffusion-based image editing via self-consistency, reporting average gains across editors and backbones.

  38. Follow-Your-Preference++: Rethinking Preference Alignment for Image Inpainting

    cs.CV 2026-06 unverdicted novelty 3.0

    Empirical study shows reward model ensembles mitigate biases like brightness and composition in preference data for image inpainting, yielding better performance than prior methods without architecture changes.