Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T05:43:00.320438Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2607.13278.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T05:43:00.320438Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T05:42:55.778574Z
A source-named dated measurement, never combined with another source.
Source: cited_works
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 822ef207-e4a6-4da2-970a-b474c3715625 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb5fb129-ac8e-4407-b9ae-82e023d955d6 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1ad2883-6a7e-49b7-91ec-280ce6d5d601 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion To convert the generated mel spectrogram into a waveform, we use an off-the-shelf BigVGAN vocoder [17]
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38e0aa8e-3051-46be-9277-36a5a710f9d2 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion completely different person,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b436b76d-67cd-4684-97ed-1f98665be0b3 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79cac068-3f77-428f-89b4-36304c6ec928 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion 500643750 (MU 2686/15-1)
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3124169-3351-413d-8735-8adf27b72eff · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Diffusion-based voice conversion with fast maximum likelihood sampling scheme,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce2583b1-9106-4b56-9ace-c8cc9f1e9830 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion DiffSinger: Singing voice synthesis via shallow diffusion mechanism,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d357ad7-6517-4d73-a4a3-be1bbd797971 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion An overview of voice conversion and its challenges: From statistical modeling to deep learning,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ce2fd71-6cdc-4853-89f0-1fd34bee60fe · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Expres- siveSinger: Multilingual and multi-style score-based singing voice synthesis with expressive performance control,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44e14be7-b8a4-4ae8-9416-8da2c86accef · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Everyone- Can-Sing: Zero-shot singing voice synthesis and conversion with speech reference,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe384600-68c8-4b51-84d5-18a79ac1b4b5 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Multi-instrument music synthesis with spectrogram diffusion,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2139c2d0-8cf6-4b8d-89d6-88db99bac30d · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Per- formance conditioning for diffusion-based multi-instrument music synthesis,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1d1a0fd-fec2-4f39-8284-b3fc423a2396 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Multi- aspect conditioning for diffusion-based music synthesis: En- hancing realism and acoustic control,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50832927-8976-4a0d-8dfe-3419521ed9a9 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion StarGAN- VC: Non-parallel many-to-many voice conversion using star generative adversarial networks,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d42b982-0b89-4b26-bd28-e7411cdab9ad · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Simple and controllable music gen- eration,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 440c842c-c683-4b6b-87a9-8bb36e92fefa · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Matcha-TTS: A fast TTS architecture with conditional flow matching,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dea6aab-b77c-4113-afc9-4ff5e9f453ad · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion FlowMac: Con- ditional flow matching for audio coding at low bit rates,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 318a15ba-d87d-430e-9436-5faabc834160 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion PAD-VC: A prosody-aware decoder for any- to-few voice conversion,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c237d231-a133-4cc0-95a3-9354be1749ad · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion High-fidelity neu- ral phonetic posteriorgrams,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cb7ec2d-9fe0-410b-a735-0a931cc6a2dd · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion SingStyle111: A multilingual singing dataset with style trans- fer,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2c0a674-e693-429d-b027-96fc885a8b93 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion NaturalSpeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cbb7213-1648-42e1-9a34-9a9891bf9155 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Hybrid transformers for music source separation,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24d416ad-4ec7-4d13-9fc2-993f50786275 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion BigVGAN: A universal neural vocoder with large-scale train- ing,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd3bd408-7748-41a0-a687-a1cea94706ed · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion FiLM: Visual reasoning with a general conditioning layer,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1559db2-605a-4496-976f-b60b999a59cb · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Classifier-Free Diffusion Guidance
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 580ab115-1d8e-42ed-b35d-89d014d2cc0f · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion CREPE: A con- volutional representation for pitch estimation,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c4afe25-a42c-4a74-b583-b12a696e6396 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion wav2vec 2.0: A framework for self-supervised learning of speech rep- resentations,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4eaf040-aed7-4ce3-8fc0-5753276e29a7 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion XLS-R: Self-supervised cross-lingual speech representation learning at scale,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01395bc0-fee1-4007-84c1-e179f7acdcc8 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Schubert Winterreise dataset: A multimodal scenario for music analysis,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd4cacfd-35e7-4c19-a9ec-32cb92157f3f · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Onsets and Frames: Dual-objective piano transcription,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a8d2515-8d64-43fb-a36b-e51a06485fcd · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Enabling factorized piano music modeling and generation with the MAESTRO dataset,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87d8f3df-0168-4f5b-b70c-af3e5c8fd6d0 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Unaligned supervision for au- tomatic music transcription in the wild,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06b02c16-231e-4477-93b9-c10d75ee119c · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Count The Notes: Histogram-based supervision for automatic music tran- scription,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b0751c0-baeb-4bb5-8d7b-823e0dc3e948 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Towards learning a universal non-semantic representation of speech,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8925945e-e0e5-447f-857b-55dee9589c88 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Bridging the training–inference gap in TTS: Training strategies for robust generative postprocessing for low-resource speakers,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7edfae29-0b21-4f0f-8098-bf076fce0c03 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Jensen–Shannon divergence and Hilbert space embedding,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c290ee8-3722-419c-b744-c0c6dd63e50d · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Fr ´echet Audio Distance: A reference-free metric for evaluating music enhancement algorithms,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 515f9d98-20ae-40a4-b573-73589c73d184 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion ITU-R Rec. BS.1534-3: Method for the subjective assessment of interme- diate quality levels of coding systems,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2001e496-a060-47f6-9528-ec8b0276afd6 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion webMUSHRA—a comprehen- sive framework for web-based listening tests,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a39fa90-9536-4978-a36b-f1ac3783d6e9 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Melody transcription from music audio: Ap- proaches and evaluation,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74c85c7e-8d73-4f46-a6a9-9d24319c01b5 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Melody extraction from poly- phonic music signals using pitch contour characteristics,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55e81e26-1c9a-484b-a99b-b405e63efeee · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion MIR EV AL: A transparent im- plementation of common MIR metrics,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59db7761-fe49-4ef4-8526-fb6d97597c6b · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Markov processes over denumerable prod- ucts of spaces, describing large systems of automata,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2b95b7c-ff9e-4100-a14d-dcee5e993cb1 · outbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion 11 020–11 028
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb5fb129-ac8e-4407-b9ae-82e023d955d6 · inbound
Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.