Pith. sign in

Paper Citation Record · LEDGER

Audio Editing in the Era of Foundation Models: A Survey

As of 13 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2606.23139.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.23139 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T07:10:49.047784Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:46.893801Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:14:51.965614Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact9
  • verified fuzzy0
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cb4f76ea-3110-4914-9eb0-bcab61f47a1d · outbound

This paper cites Beyond Voice Identity Conversion: Manipulating Voice Attributes by Adversarial Learning of Structured Disentangled Representations.

Audio Editing in the Era of Foundation Models: A Survey Beyond Voice Identity Conversion: Manipulating Voice Attributes by Adversarial Learning of Structured Disentangled Representations

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:09:49.446364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:931942a9ba316293370d92426211ca7e2c505b84b549200b09a792334ed08b28

Observation 94582d15-3150-40de-928a-1554ac99539e · outbound

This paper cites CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models.

Audio Editing in the Era of Foundation Models: A Survey CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T12:09:49.438815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:96bba7cafe489b605df2295e9b0480d4d81a3323af4bfd8fbc1b8aa7d168723a

Observation 896d5d7c-a539-41fa-b14a-32f9678cde2f · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Audio Editing in the Era of Foundation Models: A Survey Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T12:09:49.448839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:65d0797e9d3fa841fae2e8516b2cf2de358e61fb1d96b567fe0ac439e62f65f4

Observation dd9c876d-64c2-4be7-adbd-133aae9c09ee · outbound

This paper cites InForty-first interna- tional conference on machine learning.

Audio Editing in the Era of Foundation Models: A Survey InForty-first interna- tional conference on machine learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T07:10:49.047784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:08d1c9367c6753826a4ffcb578e2aace2c48b0ff130fa3386bcc6d38608557c4

Observation e4efb388-6349-4dc3-9823-e30a4318b712 · outbound

This paper cites MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement.

Audio Editing in the Era of Foundation Models: A Survey MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:09:49.443926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:5369fc635caded2867fbb8a3cf40b06dbab92bdf88463fa3b9bd77916149fc16

Observation 6ec8d7f1-8f17-47db-9598-021d3bb1e87a · outbound

This paper cites FunASR: A Fundamental End-to-End Speech Recognition Toolkit.

Audio Editing in the Era of Foundation Models: A Survey FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:09:49.451429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:1b142d193775701129d73176afe1ee82082acd1249eff6478e2fefb392743493

Observation 5bff76c7-2b3a-478d-bc3e-aaf2a390e528 · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

Audio Editing in the Era of Foundation Models: A Survey WavChat: A Survey of Spoken Dialogue Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:09:49.441309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:ee1a7ca5dae01b1ac399f089871b5550eb3062870baab35bf088bd9d6fe06183

Observation 62ddaa29-ccb2-47a4-911d-59ebd14e2bcb · outbound

This paper cites DGMO: Training-Free Audio Source Separation through Diffusion-Guided Mask Optimization.

Audio Editing in the Era of Foundation Models: A Survey DGMO: Training-Free Audio Source Separation through Diffusion-Guided Mask Optimization

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:09:49.456447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:f96528ba2511e46092a65eb0ca7a54072d4747407c6b28e7b3f6383161f0f634

Observation d12137e7-5982-4350-a79a-a27cfaf0e2fc · outbound

This paper cites AudioMorphix: Training-free audio editing with diffusion probabilistic models.

Audio Editing in the Era of Foundation Models: A Survey AudioMorphix: Training-free audio editing with diffusion probabilistic models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:09:49.461540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:98c23d1c2afd8c7c1519b396efd11b2a05bbe7c5859bc7915df534a048a77e08

Observation 5c29b7a7-62cb-4922-8fb1-85cb048405d2 · outbound

This paper cites Chen-Chou Lo, Szu-Wei Fu, Wen-Chin Huang, Xin Wang, Junichi Yamagishi, Yu Tsao, and Hsin-Min Wang.

Audio Editing in the Era of Foundation Models: A Survey Chen-Chou Lo, Szu-Wei Fu, Wen-Chin Huang, Xin Wang, Junichi Yamagishi, Yu Tsao, and Hsin-Min Wang

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:09:49.453966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:9c443adc5b92550801321f9f9e698e76feddc30f9e1aedf4c2f97ca429c06c2b

Observation 32873808-289a-4db8-8c88-9929310350bf · outbound

This paper cites Audio Editing with Non-Rigid Text Prompts.

Audio Editing in the Era of Foundation Models: A Survey Audio Editing with Non-Rigid Text Prompts

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:09:49.458732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:1c048cd7a315ecf480ffd0c5fdae24d9b5672341501d5cc05c0fc53dc60c03b7

Observation c31157f7-ed2f-4a12-9335-966e9d650b62 · outbound

This paper cites Alessandro Ragano, Jan Skoglund, and Andrew Hines.

Audio Editing in the Era of Foundation Models: A Survey Alessandro Ragano, Jan Skoglund, and Andrew Hines

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T07:10:49.047784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:ed54b2393d19fec39e015bda950ded3a71e48a2ef40ee35b5b78b481fcdb08be

Observation 99705ab8-2efa-47e9-a1bd-1a5a30aa568e · outbound

This paper cites InICASSP 2024-2024 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP), pages 1011–1015.

Audio Editing in the Era of Foundation Models: A Survey InICASSP 2024-2024 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP), pages 1011–1015

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T07:10:49.047784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:8043a256bf2f6ac1d9a0430138795ea7008a089e6eb5d8ff3331aaf4ff8591eb

Observation c5fa8f1f-5d05-4aa1-8956-3fbde8d9f65b · outbound

This paper cites Seedance 2.0: Advancing Video Generation for World Complexity.

Audio Editing in the Era of Foundation Models: A Survey Seedance 2.0: Advancing Video Generation for World Complexity

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T12:09:49.466189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:a071edd5fb3d37e6764d01d02bd48e8417790b43bed7ccd5a1c2860999f8c460

Observation 7824c1a5-e363-4d2e-96a7-6333d233b71b · outbound

This paper cites EdiTTS: Score-based Editing for Controllable Text-to-Speech.

Audio Editing in the Era of Foundation Models: A Survey EdiTTS: Score-based Editing for Controllable Text-to-Speech

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:09:49.463871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:1c41ee6c5de58b45c00599c1865a382481c33853c1506d2b10bd37dd10026b9e

Observation 71f28937-d60a-405b-a0be-14662efa2971 · outbound

This paper cites RMVPE: A Robust Model for Vocal Pitch Estimation in Polyphonic Music.

Audio Editing in the Era of Foundation Models: A Survey RMVPE: A Robust Model for Vocal Pitch Estimation in Polyphonic Music

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:09:49.471000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:80429beb06bac570887d3e8c47f159a59d127000f9577a058bd10c7993b57ba9

Observation 79789653-23c4-4158-b840-b3078efe1965 · outbound

This paper cites Qwen3-Omni Technical Report.

Audio Editing in the Era of Foundation Models: A Survey Qwen3-Omni Technical Report

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T12:09:49.468540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:d937eaba080ddb1e38fa7956eedc50ee7fc98cafbf819bdd6586176544cc0488

Observation 9b5921d0-3a44-4430-ab2a-1aa7512fca92 · outbound

This paper cites Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction Tuning.

Audio Editing in the Era of Foundation Models: A Survey Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction Tuning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:09:49.473578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:4f031f27e8d3d68bc21b9e928e0231f47eb6630c587efd8a51a78fd2dd6835ba

Observation 2135d651-c307-4c94-b3ec-073964a0d8c0 · outbound

This paper cites When the editable unit is defined by speaker activity rather than text, pyannote (Bredin,.

Audio Editing in the Era of Foundation Models: A Survey When the editable unit is defined by speaker activity rather than text, pyannote (Bredin,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T07:10:49.047784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:8ed020735835f51271d1747c654a84f67299aadbfae95be93b69f0fc68f3f269

Observation b86f0995-a9a7-4ee3-a6ec-d00deea93a91 · outbound

This paper cites For general au- dio, sound event detection models (Kong et al., 2020; Li et al., 2023) produce event-level activity boundaries.

Audio Editing in the Era of Foundation Models: A Survey For general au- dio, sound event detection models (Kong et al., 2020; Li et al., 2023) produce event-level activity boundaries

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T07:10:49.047784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:aa33917ced37b82c82d54846332e79281ce9349e176182436dd4a1c6f5c4f766

Observation b3cb866c-a75c-409b-84f4-59ba612d096c · outbound

This paper cites How- ever, their effectiveness depends heavily on stable text-acoustic alignment and high-quality tokeniza- tion.

Audio Editing in the Era of Foundation Models: A Survey How- ever, their effectiveness depends heavily on stable text-acoustic alignment and high-quality tokeniza- tion

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-26T07:10:49.047784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:5a2fca5bd67b66f69ac001a91a63aeeeacdecf211f01a9d02ef8316a184ba8ee

Pith citing papers

Observation 63e73bf0-5802-41dd-a8da-044596636ada · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Audio Editing in the Era of Foundation Models: A Survey

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:26.450915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:26.450915Z digest=sha256:c9b9d2fc1c5357f4a9df7fdcbe698aef30b1ab691b06424b952b2c296f3c09aa

Observation 131edcb5-940e-44b9-96ce-f9d35819cbee · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Audio Editing in the Era of Foundation Models: A Survey

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:14:52.074519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T00:14:46.893801Z digest=sha256:7615703b49f13dbcd117fb1c9d7c170805aed669f3a2507487cdc05711363efd