Pith. sign in

REVIEW 2 cited by

Prompt-guided Precise Audio Editing with Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.04350 v1 pith:NCQE2A5G submitted 2024-05-11 cs.SD cs.AIcs.LGeess.AS

classification cs.SDcs.AIcs.LGeess.AS
keywords editingaudiodiffusionmodelspreciseaccurateadvancementsalthough
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Audio editing involves the arbitrary manipulation of audio content through precise control. Although text-guided diffusion models have made significant advancements in text-to-audio generation, they still face challenges in finding a flexible and precise way to modify target events within an audio track. We present a novel approach, referred to as PPAE, which serves as a general module for diffusion models and enables precise audio editing. The editing is based on the input textual prompt only and is entirely training-free. We exploit the cross-attention maps of diffusion models to facilitate accurate local editing and employ a hierarchical local-global pipeline to ensure a smoother editing process. Experimental results highlight the effectiveness of our method in various editing tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. One Prompt, Many Sounds: Modeling Listener Variability in LLM-Based Equalization

    cs.SD 2026-01 unverdicted novelty 6.0 of 10

    LLMs using in-context learning and fine-tuning on listener experiment data generate equalization settings that align better with population preferences than random sampling or static presets.

  2. DGMO: Training-Free Audio Source Separation through Diffusion-Guided Mask Optimization

    eess.AS 2025-06 conditional novelty 6.0 of 10

    Diffusion-Guided Mask Optimization shows a frozen text-to-audio diffusion model can perform zero-shot language-queried source separation by fitting a spectrogram mask to the model's generated reference.

Pith tools