Pith. sign in

REVIEW 5 cited by

FluxSpace: Disentangled Semantic Editing in Rectified Flow Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.09611 v1 pith:QULGJFIO submitted 2024-12-12 cs.CV

classification cs.CV
keywords imageeditingflowrectifiedmodelsabilityapproachcapabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Rectified flow models have emerged as a dominant approach in image generation, showcasing impressive capabilities in high-quality image synthesis. However, despite their effectiveness in visual generation, rectified flow models often struggle with disentangled editing of images. This limitation prevents the ability to perform precise, attribute-specific modifications without affecting unrelated aspects of the image. In this paper, we introduce FluxSpace, a domain-agnostic image editing method leveraging a representation space with the ability to control the semantics of images generated by rectified flow transformers, such as Flux. By leveraging the representations learned by the transformer blocks within the rectified flow models, we propose a set of semantically interpretable representations that enable a wide range of image editing tasks, from fine-grained image editing to artistic creation. This work offers a scalable and effective image editing approach, along with its disentanglement capabilities.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Flow Straight and Fast in Hilbert Space: Functional Rectified Flow

    cs.LG 2025-09 conditional novelty 7.0 of 10

    Functional rectified flow is defined and proved to preserve marginals in separable Hilbert spaces, with functional flow matching and probability-flow ODEs as special cases.

  2. SEED: A Benchmark Dataset for Sequential Facial Attribute Editing with Diffusion Models

    cs.CV 2025-05 reject novelty 6.0 of 10

    SEED is a 91,526-image benchmark of diffusion-generated sequential facial edits with sequence, mask, and prompt annotations, and FAITH adds DWT high-frequency cues to a transformer for edit-sequence detection.

  3. WordCon: Word-level Typography Control in Scene Text Rendering

    cs.CV 2025-06 conditional novelty 5.0 of 10

    WordCon uses grounding-model masks and two extra losses to fine-tune Flux so that typography can be controlled word by word.

  4. PairEdit: Learning Semantic Variations for Exemplar-based Image Editing

    cs.CV 2025-06 conditional novelty 5.0 of 10

    PairEdit trains two LoRA adapters on a pretrained diffusion model to capture the semantic direction between paired source-target images, enabling text-free, controllable image editing from as few as one pair.

  5. DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing

    cs.CV 2025-06 conditional novelty 4.0 of 10

    DFVEdit edits videos by iteratively subtracting a conditional delta flow vector, the difference between the model's predictions under the target and source prompts, from the latent representation of the source video.

Pith tools