Pith. sign in

Paper Citation Record · LEDGER

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion

As of 20 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2607.13278.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13278 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:43:00.320438Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:42:55.778574Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 822ef207-e4a6-4da2-970a-b474c3715625 · outbound

This paper cites an unresolved cited work.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:55.624158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:55.624158Z digest=sha256:e716b96bd04d920f806ccbaf54b7b93e85a54cd7aef380d1e519255cfb6df456

Observation eb5fb129-ac8e-4407-b9ae-82e023d955d6 · outbound

This paper cites Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:55.778574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:55.778574Z digest=sha256:76edd020228b312010ac3b49180c93ec96e07edfd77892d4f074e98ebb0c8d24

Observation b1ad2883-6a7e-49b7-91ec-280ce6d5d601 · outbound

This paper cites To convert the generated mel spectrogram into a waveform, we use an off-the-shelf BigVGAN vocoder [17].

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion To convert the generated mel spectrogram into a waveform, we use an off-the-shelf BigVGAN vocoder [17]

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:55.926675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:55.926675Z digest=sha256:af7deee74fe572a9f26319774aedc78c6f4bbc0ebc673aec71f9fd9b3b53d830

Observation 38e0aa8e-3051-46be-9277-36a5a710f9d2 · outbound

This paper cites completely different person,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion completely different person,

Reference 4

Resolution
malformed identifier
no resolver link, observed 2026-08-02T05:42:56.075313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:56.075313Z digest=sha256:934673eb21bdea695b607f4de21be836fd637e8f58df3964ddab72027be3b1e3

Observation b436b76d-67cd-4684-97ed-1f98665be0b3 · outbound

This paper cites an unresolved cited work.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:56.232755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:56.232755Z digest=sha256:4da12b3aaebb0e350541f072ebe32d432fa64968f2b0a45bd0bef648075d2fb7

Observation 79cac068-3f77-428f-89b4-36304c6ec928 · outbound

This paper cites 500643750 (MU 2686/15-1).

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion 500643750 (MU 2686/15-1)

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:56.428388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:56.428388Z digest=sha256:c87b93a05ac3488d167d3782e3628d138753fb3014be79e00fc2f73111f01651

Observation a3124169-3351-413d-8735-8adf27b72eff · outbound

This paper cites Diffusion-based voice conversion with fast maximum likelihood sampling scheme,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Diffusion-based voice conversion with fast maximum likelihood sampling scheme,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:56.554187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:56.554187Z digest=sha256:c8c8d21740a2518044ff0cc8f05d47fe912646c6021faec519d34c3bd48c8b55

Observation ce2583b1-9106-4b56-9ace-c8cc9f1e9830 · outbound

This paper cites DiffSinger: Singing voice synthesis via shallow diffusion mechanism,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion DiffSinger: Singing voice synthesis via shallow diffusion mechanism,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:56.647106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:56.647106Z digest=sha256:0b7da92d1ef28ee8a56998a1dda32754854c89046e02153868f8373243f2519a

Observation 2d357ad7-6517-4d73-a4a3-be1bbd797971 · outbound

This paper cites An overview of voice conversion and its challenges: From statistical modeling to deep learning,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion An overview of voice conversion and its challenges: From statistical modeling to deep learning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:57.697864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:57.697864Z digest=sha256:d5392eb0cd489bc1635735e0ac85bd0cceb282be0414d8bfaedee45d10139f62

Observation 2ce2fd71-6cdc-4853-89f0-1fd34bee60fe · outbound

This paper cites Expres- siveSinger: Multilingual and multi-style score-based singing voice synthesis with expressive performance control,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Expres- siveSinger: Multilingual and multi-style score-based singing voice synthesis with expressive performance control,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:56.872722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:56.872722Z digest=sha256:fc95ed93b59e7724ec72415b157b0816eb6dadf887d4b04afe5b5271307a4866

Observation 44e14be7-b8a4-4ae8-9416-8da2c86accef · outbound

This paper cites Everyone- Can-Sing: Zero-shot singing voice synthesis and conversion with speech reference,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Everyone- Can-Sing: Zero-shot singing voice synthesis and conversion with speech reference,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:56.970918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:56.970918Z digest=sha256:75fc53c8b6250053bef79a716f1f7a9f6c3bac6cacdf53173a88342981e16676

Observation fe384600-68c8-4b51-84d5-18a79ac1b4b5 · outbound

This paper cites Multi-instrument music synthesis with spectrogram diffusion,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Multi-instrument music synthesis with spectrogram diffusion,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:57.072171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:57.072171Z digest=sha256:181e954ec5e08514a733003824c1f316fc95bac703076a758a7a29f130c8d924

Observation 2139c2d0-8cf6-4b8d-89d6-88db99bac30d · outbound

This paper cites Per- formance conditioning for diffusion-based multi-instrument music synthesis,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Per- formance conditioning for diffusion-based multi-instrument music synthesis,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:57.167511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:57.167511Z digest=sha256:b5707c5597e389742becd59ca1249ca408a7c1d478535ead211da4e3ece2b53d

Observation d1d1a0fd-fec2-4f39-8284-b3fc423a2396 · outbound

This paper cites Multi- aspect conditioning for diffusion-based music synthesis: En- hancing realism and acoustic control,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Multi- aspect conditioning for diffusion-based music synthesis: En- hancing realism and acoustic control,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:57.319193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:57.319193Z digest=sha256:15886097dbb32cef332d043218ab6b4574961bafd40db3a6a7c67825ba778380

Observation 50832927-8976-4a0d-8dfe-3419521ed9a9 · outbound

This paper cites StarGAN- VC: Non-parallel many-to-many voice conversion using star generative adversarial networks,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion StarGAN- VC: Non-parallel many-to-many voice conversion using star generative adversarial networks,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:57.520009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:57.520009Z digest=sha256:92307f463a2bcb59cc9dfe37aab97c61e15ee1fe215ea1ab7aab9424e63babbf

Observation 2d42b982-0b89-4b26-bd28-e7411cdab9ad · outbound

This paper cites Simple and controllable music gen- eration,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Simple and controllable music gen- eration,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:58.456060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:58.456060Z digest=sha256:4a791f41924504a6cd20d00cd9cb6505cea5f5d4400599a6317a60d014ac9cd0

Observation 440c842c-c683-4b6b-87a9-8bb36e92fefa · outbound

This paper cites Matcha-TTS: A fast TTS architecture with conditional flow matching,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Matcha-TTS: A fast TTS architecture with conditional flow matching,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:57.810125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:57.810125Z digest=sha256:269e15af953e0a991de4e036fcb09d14e7b87c8c82c316e7263fe68990074093

Observation 5dea6aab-b77c-4113-afc9-4ff5e9f453ad · outbound

This paper cites FlowMac: Con- ditional flow matching for audio coding at low bit rates,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion FlowMac: Con- ditional flow matching for audio coding at low bit rates,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:57.912421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:57.912421Z digest=sha256:280846d20595a3fc4cf0291364f7d6613b22776180354fa8af1735a32674426c

Observation 318a15ba-d87d-430e-9436-5faabc834160 · outbound

This paper cites PAD-VC: A prosody-aware decoder for any- to-few voice conversion,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion PAD-VC: A prosody-aware decoder for any- to-few voice conversion,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:58.015661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:58.015661Z digest=sha256:5716c18bcfc270671c2487fb962dd914474d87a1dab32f11a2cbe488a1c5b74e

Observation c237d231-a133-4cc0-95a3-9354be1749ad · outbound

This paper cites High-fidelity neu- ral phonetic posteriorgrams,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion High-fidelity neu- ral phonetic posteriorgrams,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:58.132976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:58.132976Z digest=sha256:ae2235042d02e749d975e5d92825e913d1c03850d65435cc65f6a1d119f1f4dc

Observation 1cb7ec2d-9fe0-410b-a735-0a931cc6a2dd · outbound

This paper cites SingStyle111: A multilingual singing dataset with style trans- fer,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion SingStyle111: A multilingual singing dataset with style trans- fer,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:58.276776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:58.276776Z digest=sha256:e9a06f16a420ae880fff86e772735795d1e77945b921b33487ae9edee6327fc0

Observation b2c0a674-e693-429d-b027-96fc885a8b93 · outbound

This paper cites NaturalSpeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion NaturalSpeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:58.375252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:58.375252Z digest=sha256:95b482803567b572dfb7198600f74f49bcb00112de831e799630a4f1d14c6a5d

Observation 7cbb7213-1648-42e1-9a34-9a9891bf9155 · outbound

This paper cites Hybrid transformers for music source separation,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Hybrid transformers for music source separation,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:58.866612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:58.866612Z digest=sha256:fe4d407bc16abc50cf9a3361c2fb6831601cbb17803783b2726c5db10c7f8ad5

Observation 24d416ad-4ec7-4d13-9fc2-993f50786275 · outbound

This paper cites BigVGAN: A universal neural vocoder with large-scale train- ing,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion BigVGAN: A universal neural vocoder with large-scale train- ing,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:58.539092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:58.539092Z digest=sha256:67930cd8228d8c8d68bee0819b802600f01463f0ca17334d523def4797324f35

Observation cd3bd408-7748-41a0-a687-a1cea94706ed · outbound

This paper cites FiLM: Visual reasoning with a general conditioning layer,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion FiLM: Visual reasoning with a general conditioning layer,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:58.597933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:58.597933Z digest=sha256:d2a6e6ec30bd947ed7cad455a0b9fedca842ca8f2dc144dc05776cd77483ac72

Observation f1559db2-605a-4496-976f-b60b999a59cb · outbound

This paper cites Classifier-Free Diffusion Guidance.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Classifier-Free Diffusion Guidance

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:58.648482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:58.648482Z digest=sha256:dd23db89b6441ff03fe14b34bce3c841d7b4846d302d77488cde22ab1235a670

Observation 580ab115-1d8e-42ed-b35d-89d014d2cc0f · outbound

This paper cites CREPE: A con- volutional representation for pitch estimation,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion CREPE: A con- volutional representation for pitch estimation,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:58.732592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:58.732592Z digest=sha256:57cfb78f055fd3483dbd1115bb2f2df5f85c9fb6e42c55dacc990c1428f0673e

Observation 0c4afe25-a42c-4a74-b583-b12a696e6396 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech rep- resentations,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion wav2vec 2.0: A framework for self-supervised learning of speech rep- resentations,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:58.791069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:58.791069Z digest=sha256:78de5fc200cd31e204c1cd82572c119ee43770ef09f64c23e0f1c05ab8a4921b

Observation d4eaf040-aed7-4ce3-8fc0-5753276e29a7 · outbound

This paper cites XLS-R: Self-supervised cross-lingual speech representation learning at scale,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion XLS-R: Self-supervised cross-lingual speech representation learning at scale,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:58.862919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:58.862919Z digest=sha256:ffcefd9556819ef91a82057e14924e3035d94e2e2592ab26ca1b17c0f64583e5

Observation 01395bc0-fee1-4007-84c1-e179f7acdcc8 · outbound

This paper cites Schubert Winterreise dataset: A multimodal scenario for music analysis,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Schubert Winterreise dataset: A multimodal scenario for music analysis,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:59.612229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:59.612229Z digest=sha256:ac46e77e7193fad500d5bdffaefe0deea071bf6b3154a22e80fde5d5f611f80f

Observation cd4cacfd-35e7-4c19-a9ec-32cb92157f3f · outbound

This paper cites Onsets and Frames: Dual-objective piano transcription,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Onsets and Frames: Dual-objective piano transcription,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:58.891463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:58.891463Z digest=sha256:4599ea2f17496e3afdb63b3b3ba49b478f36e5064c191c528972cc246df90691

Observation 0a8d2515-8d64-43fb-a36b-e51a06485fcd · outbound

This paper cites Enabling factorized piano music modeling and generation with the MAESTRO dataset,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Enabling factorized piano music modeling and generation with the MAESTRO dataset,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:58.926796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:58.926796Z digest=sha256:68d1f0de0739ea636c89987ef64035b5288db146d871397b3d3fc5fd26f1031d

Observation 87d8f3df-0168-4f5b-b70c-af3e5c8fd6d0 · outbound

This paper cites Unaligned supervision for au- tomatic music transcription in the wild,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Unaligned supervision for au- tomatic music transcription in the wild,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:59.044184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:59.044184Z digest=sha256:7b370e10cd8aa265a8d2c91571fef3451cb5e6b7bd8c468f74dba176c5e37993

Observation 06b02c16-231e-4477-93b9-c10d75ee119c · outbound

This paper cites Count The Notes: Histogram-based supervision for automatic music tran- scription,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Count The Notes: Histogram-based supervision for automatic music tran- scription,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:59.145050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:59.145050Z digest=sha256:0de8f95530dfe5d090bde4560752b72bdbad8066a988be26cdf953adc659f885

Observation 8b0751c0-baeb-4bb5-8d7b-823e0dc3e948 · outbound

This paper cites Towards learning a universal non-semantic representation of speech,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Towards learning a universal non-semantic representation of speech,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:59.297638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:59.297638Z digest=sha256:6db18d0b14b76ecb981b081c4fc2b8d26db478330c5afabbd5e3bf6bfeb96338

Observation 8925945e-e0e5-447f-857b-55dee9589c88 · outbound

This paper cites Bridging the training–inference gap in TTS: Training strategies for robust generative postprocessing for low-resource speakers,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Bridging the training–inference gap in TTS: Training strategies for robust generative postprocessing for low-resource speakers,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:59.456524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:59.456524Z digest=sha256:064bf3e5093498fb2c785dc483b3db105f225feac3082557bcd802fca5f1c77c

Observation 7edfae29-0b21-4f0f-8098-bf076fce0c03 · outbound

This paper cites Jensen–Shannon divergence and Hilbert space embedding,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Jensen–Shannon divergence and Hilbert space embedding,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T05:43:00.285800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:43:00.285800Z digest=sha256:05b48affc00dad9b10f59fa24cac36f89f890cfa315cecca62fcead5f97aa1d6

Observation 7c290ee8-3722-419c-b744-c0c6dd63e50d · outbound

This paper cites Fr ´echet Audio Distance: A reference-free metric for evaluating music enhancement algorithms,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Fr ´echet Audio Distance: A reference-free metric for evaluating music enhancement algorithms,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:59.780120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:59.780120Z digest=sha256:b631493e73616cb9cff96b44ec7c0cf70f3e126a1cba1fa302f1ce907977bc9c

Observation 515f9d98-20ae-40a4-b573-73589c73d184 · outbound

This paper cites ITU-R Rec. BS.1534-3: Method for the subjective assessment of interme- diate quality levels of coding systems,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion ITU-R Rec. BS.1534-3: Method for the subjective assessment of interme- diate quality levels of coding systems,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:59.892013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:59.892013Z digest=sha256:a2d939efb8e1cfe062204136b907002884ac4cff36a5cbace5d24e3a53a06cda

Observation 2001e496-a060-47f6-9528-ec8b0276afd6 · outbound

This paper cites webMUSHRA—a comprehen- sive framework for web-based listening tests,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion webMUSHRA—a comprehen- sive framework for web-based listening tests,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:59.902579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:59.902579Z digest=sha256:76c5ad5837d4d8342bcb71750422cc207ac2bf51c9960a1ce12bc545eaf55884

Observation 9a39fa90-9536-4978-a36b-f1ac3783d6e9 · outbound

This paper cites Melody transcription from music audio: Ap- proaches and evaluation,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Melody transcription from music audio: Ap- proaches and evaluation,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:59.982555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:59.982555Z digest=sha256:5559284e9c16dd32650226c03d35a0dce1ebc74656b291a1715ed1ece6ffa706

Observation 74c85c7e-8d73-4f46-a6a9-9d24319c01b5 · outbound

This paper cites Melody extraction from poly- phonic music signals using pitch contour characteristics,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Melody extraction from poly- phonic music signals using pitch contour characteristics,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T05:43:00.083433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:43:00.083433Z digest=sha256:075093fbad20fd54d596de653e05c80455b3edffef715a1132ab8c3d70c58a3e

Observation 55e81e26-1c9a-484b-a99b-b405e63efeee · outbound

This paper cites MIR EV AL: A transparent im- plementation of common MIR metrics,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion MIR EV AL: A transparent im- plementation of common MIR metrics,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T05:43:00.184615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:43:00.184615Z digest=sha256:bc05fadfb8c87d11fe06acd807810419732c86635d78537c8a269cfa3f8ea2a0

Observation 59db7761-fe49-4ef4-8526-fb6d97597c6b · outbound

This paper cites Markov processes over denumerable prod- ucts of spaces, describing large systems of automata,.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Markov processes over denumerable prod- ucts of spaces, describing large systems of automata,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T05:43:00.320438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:43:00.320438Z digest=sha256:8bc0d8e75997c70177b6fe0f9b5771e68b686b25bd8670a6ff5a5a94220f3493

Observation e2b95b7c-ff9e-4100-a14d-dcee5e993cb1 · outbound

This paper cites 11 020–11 028.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion 11 020–11 028

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:56.746440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:56.746440Z digest=sha256:201684c67b44fcb3d8003e6fa37571eefee1a4ecc0fcf86d5a4d60606ff65700

Pith citing papers

Observation eb5fb129-ac8e-4407-b9ae-82e023d955d6 · inbound

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion cites this paper.

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T05:42:55.778574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:42:55.778574Z digest=sha256:76edd020228b312010ac3b49180c93ec96e07edfd77892d4f074e98ebb0c8d24