Pith. sign in

REVIEW 7 cited by

AudioMorphix: Training-free audio editing with diffusion probabilistic models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.16076 v1 pith:PIVOGWKQ submitted 2025-05-21 eess.AS

classification eess.AS
keywords audioeditingaudiomorphixwhilecontentdiffusionfidelityinstructions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Editing sound with precision is a crucial yet underexplored challenge in audio content creation. While existing works can manipulate sounds by text instructions or audio exemplar pairs, they often struggled to modify audio content precisely while preserving fidelity to the original recording. In this work, we introduce a novel editing approach that enables localized modifications to specific time-frequency regions while keeping the remaining of the audio intact by operating on spectrograms directly. To achieve this, we propose AudioMorphix, a training-free audio editor that manipulates a target region on the spectrogram by referring to another recording. Inspired by morphing theory, we conceptualize audio mixing as a process where different sounds blend seamlessly through morphing and can be decomposed back into individual components via demorphing. Our AudioMorphix optimizes the noised latent conditioned on raw input and reference audio while rectifying the guided diffusion process through a series of energy functions. Additionally, we enhance self-attention layers with a cache mechanism to preserve detailed characteristics from the original recordings. To advance audio editing research, we devise a new evaluation benchmark, which includes a curated dataset with a variety of editing instructions. Extensive experiments demonstrate that AudioMorphix yields promising performance on various audio editing tasks, including addition, removal, time shifting and stretching, and pitch shifting, achieving high fidelity and precision. Demo and code are available at this url.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. MMAE: A Massive Multitask Audio Editing Benchmark

    cs.SD 2026-06 conditional novelty 8.0 of 10

    MMAE is a new multitask audio editing benchmark showing that leading models achieve under 5% exact match rate, with 0% on complex mixed-modality tasks.

  2. RIME: Enabling Large-Scale Agentic Music Post-Production

    cs.SD 2026-07 conditional novelty 6.0 of 10

    RIME generates 3,000 synthetic music post-production edit triples and shows that current multimodal LLM agents can recover edit structure but often fail to set effect parameters correctly.

  3. RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers

    cs.SD 2026-06 unverdicted novelty 6.0 of 10

    Hybrid two-stage diffusion transformer architecture for instruction-guided audio editing via rectified flow that performs joint attention at low resolution then alternates joint and cross-attention at high resolution ...

  4. SemanticAudio: Audio Generation and Editing in Semantic Space

    eess.AS 2026-01 conditional novelty 6.0 of 10

    SemanticAudio improves text-to-audio alignment by generating a compact semantic plan first with a Flow Matching planner and then rendering acoustic latents from that plan, and it performs training-free audio editing b...

  5. RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers

    cs.SD 2026-06 conditional novelty 5.0 of 10

    A coarse-to-fine hybrid MMDiT/DiT audio editor trained with rectified flow matching improves fidelity and cuts edit time versus prior instruction-guided baselines on synthetic overlapping-event tasks.

  6. RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers

    cs.SD 2026-06 unverdicted novelty 5.0 of 10

    Hybrid two-stage diffusion transformer architecture for instruction-guided audio editing uses coarse joint attention at low resolution and refined alternating blocks at high resolution to improve performance and effic...

  7. Audio Editing in the Era of Foundation Models: A Survey

    eess.AS 2026-06 unverdicted novelty 3.0 of 10

    A survey that presents a unified taxonomy of audio editing tasks, summarizes training-based and training-free foundation model approaches, reviews datasets and evaluation protocols, and identifies future challenges.

Pith tools