Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T07:10:49.047784Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2606.23139.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T07:10:49.047784Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:46.893801Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T00:14:51.965614Z
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cb4f76ea-3110-4914-9eb0-bcab61f47a1d · outbound
Audio Editing in the Era of Foundation Models: A Survey Beyond Voice Identity Conversion: Manipulating Voice Attributes by Adversarial Learning of Structured Disentangled Representations
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 94582d15-3150-40de-928a-1554ac99539e · outbound
Audio Editing in the Era of Foundation Models: A Survey CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 896d5d7c-a539-41fa-b14a-32f9678cde2f · outbound
Audio Editing in the Era of Foundation Models: A Survey Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dd9c876d-64c2-4be7-adbd-133aae9c09ee · outbound
Audio Editing in the Era of Foundation Models: A Survey InForty-first interna- tional conference on machine learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4efb388-6349-4dc3-9823-e30a4318b712 · outbound
Audio Editing in the Era of Foundation Models: A Survey MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6ec8d7f1-8f17-47db-9598-021d3bb1e87a · outbound
Audio Editing in the Era of Foundation Models: A Survey FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5bff76c7-2b3a-478d-bc3e-aaf2a390e528 · outbound
Audio Editing in the Era of Foundation Models: A Survey WavChat: A Survey of Spoken Dialogue Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 62ddaa29-ccb2-47a4-911d-59ebd14e2bcb · outbound
Audio Editing in the Era of Foundation Models: A Survey DGMO: Training-Free Audio Source Separation through Diffusion-Guided Mask Optimization
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d12137e7-5982-4350-a79a-a27cfaf0e2fc · outbound
Audio Editing in the Era of Foundation Models: A Survey AudioMorphix: Training-free audio editing with diffusion probabilistic models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5c29b7a7-62cb-4922-8fb1-85cb048405d2 · outbound
Audio Editing in the Era of Foundation Models: A Survey Chen-Chou Lo, Szu-Wei Fu, Wen-Chin Huang, Xin Wang, Junichi Yamagishi, Yu Tsao, and Hsin-Min Wang
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 32873808-289a-4db8-8c88-9929310350bf · outbound
Audio Editing in the Era of Foundation Models: A Survey Audio Editing with Non-Rigid Text Prompts
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c31157f7-ed2f-4a12-9335-966e9d650b62 · outbound
Audio Editing in the Era of Foundation Models: A Survey Alessandro Ragano, Jan Skoglund, and Andrew Hines
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99705ab8-2efa-47e9-a1bd-1a5a30aa568e · outbound
Audio Editing in the Era of Foundation Models: A Survey InICASSP 2024-2024 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP), pages 1011–1015
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5fa8f1f-5d05-4aa1-8956-3fbde8d9f65b · outbound
Audio Editing in the Era of Foundation Models: A Survey Seedance 2.0: Advancing Video Generation for World Complexity
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7824c1a5-e363-4d2e-96a7-6333d233b71b · outbound
Audio Editing in the Era of Foundation Models: A Survey EdiTTS: Score-based Editing for Controllable Text-to-Speech
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 71f28937-d60a-405b-a0be-14662efa2971 · outbound
Audio Editing in the Era of Foundation Models: A Survey RMVPE: A Robust Model for Vocal Pitch Estimation in Polyphonic Music
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 79789653-23c4-4158-b840-b3078efe1965 · outbound
Audio Editing in the Era of Foundation Models: A Survey Qwen3-Omni Technical Report
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9b5921d0-3a44-4430-ab2a-1aa7512fca92 · outbound
Audio Editing in the Era of Foundation Models: A Survey Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction Tuning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2135d651-c307-4c94-b3ec-073964a0d8c0 · outbound
Audio Editing in the Era of Foundation Models: A Survey When the editable unit is defined by speaker activity rather than text, pyannote (Bredin,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b86f0995-a9a7-4ee3-a6ec-d00deea93a91 · outbound
Audio Editing in the Era of Foundation Models: A Survey For general au- dio, sound event detection models (Kong et al., 2020; Li et al., 2023) produce event-level activity boundaries
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3cb866c-a75c-409b-84f4-59ba612d096c · outbound
Audio Editing in the Era of Foundation Models: A Survey How- ever, their effectiveness depends heavily on stable text-acoustic alignment and high-quality tokeniza- tion
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63e73bf0-5802-41dd-a8da-044596636ada · inbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Audio Editing in the Era of Foundation Models: A Survey
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 131edcb5-940e-44b9-96ce-f9d35819cbee · inbound
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Audio Editing in the Era of Foundation Models: A Survey
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.