Pith. sign in

Paper Citation Record · LEDGER

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference

As of 16 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2501.13870.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13870 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:33:04.909705Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f7e4bbcc-7c6e-41c5-9c77-639c93fb05cc · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference Fastspeech 2: Fast and high-quality end-to-end text to speech

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.211518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.813257Z digest=sha256:df082f8d6bcc8f4d534af1c721a2d386217475fafd27204f23a1a2da6af13519

Observation f992f40e-592b-40c2-8010-d9e9cf3ddb1a · outbound

This paper cites Hifi- gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference Hifi- gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.201304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.818527Z digest=sha256:ff9eff5f0451949f306747a6a0b6455b99b8da84dbbaafbb79995fd64910d818

Observation 4c1bf083-b4ec-4bd0-a359-b1ddf3f32878 · outbound

This paper cites Bigvgan: A universal neural vocoder with large-scale training.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference Bigvgan: A universal neural vocoder with large-scale training

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.191035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.822366Z digest=sha256:d15402e9ada787e3c34db556bcc433c6476766f740f4f43a588a2cdce0059e16

Observation 4f293436-fa58-4733-b99f-df68a8e4b2df · outbound

This paper cites High-fidelity audio compression with improved rvqgan.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference High-fidelity audio compression with improved rvqgan

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.181645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.826444Z digest=sha256:3f31812a8df34c9180d0503586b33b114f8db8193578850d0256a7d6a76accff

Observation ae7545fd-5955-475b-9d52-dc36c4478ec0 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference Soundstream: An end-to-end neural audio codec

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.170977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.830667Z digest=sha256:7891adbbc93c1414d656601e79777b1aa5addb9938a0af1adac160854849682a

Observation e1d70811-52cc-4d14-ae3c-30c8ba576ab6 · outbound

This paper cites RMSSinger: Realistic-music-score based singing voice synthesis.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference RMSSinger: Realistic-music-score based singing voice synthesis

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.159487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.834857Z digest=sha256:bfdc0982ed77a0b933f8edd31c15074603ec1a96a1e02ccbfbefab4a0c22c68c

Observation 73f25841-d4b8-4df4-8ae0-ef2e28f59783 · outbound

This paper cites XiaoiceSing: A High-Quality and Integrated Singing Voice Synthesis System.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference XiaoiceSing: A High-Quality and Integrated Singing Voice Synthesis System

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T15:33:04.840099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:33:04.840099Z digest=sha256:14759ad96c5333ff2a1e22c7799dcb78fbf1a3993384ee9e6c3ce25ab67da77f

Observation 3d24e77e-e1d5-4003-886b-d9e53650b675 · outbound

This paper cites The singing voice con- version challenge 2023.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference The singing voice con- version challenge 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.147741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.844337Z digest=sha256:1028c4640a3d78c20e1cde826ea4b28cf42e68311b7d7d072af191db4e81379e

Observation db7b37fb-085e-4564-91a0-0beeff7de495 · outbound

This paper cites Gr0: Self-supervised global representation learning for zero-shot voice conversion.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference Gr0: Self-supervised global representation learning for zero-shot voice conversion

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.132567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.847862Z digest=sha256:131d679d48b7dfa1cfbe8bb3d491a3fe0f90cb3457d24f5a8d233c2a95631aaa

Observation f7ee6814-b9bb-4f6d-95fe-f69b508e0d05 · outbound

This paper cites Neural analysis and syn- thesis: Reconstructing speech from self-supervised rep- resentations.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference Neural analysis and syn- thesis: Reconstructing speech from self-supervised rep- resentations

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.120872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.851428Z digest=sha256:db92f6d5ec17069093b0e9b869f5261162a2c4f4f2d1db1efc57d8c2c760e616

Observation 8941964d-c425-4a6b-bd44-3b23717c4634 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T15:33:04.855210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:33:04.855210Z digest=sha256:02f49067ab67527be653241f1562005aefcba8c9b850bcbd4f95fec7ecea79f0

Observation e8e2ad1d-f0e8-43a2-84f6-02a8f55d8900 · outbound

This paper cites NANSY++: Unified Voice Synthesis with Neural Analysis and Synthesis.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference NANSY++: Unified Voice Synthesis with Neural Analysis and Synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T15:33:04.860035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:33:04.860035Z digest=sha256:c1194a2f06f8c1a7556d84e6aebf2339cbfa9dbab577d6c4b1912268695bf8c3

Observation bf479d61-27f8-4ea2-aeeb-c7097a6bbc52 · outbound

This paper cites https://github.com/svc-develop-team/so-vits-svc.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference https://github.com/svc-develop-team/so-vits-svc

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.110350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.863805Z digest=sha256:eccec01968d99ebe044f9bde08cf04f68d8572f058629b35e22c40a7c57ca2bf

Observation d92130f6-0d7a-491d-90c8-68c397ed9c31 · outbound

This paper cites Midi-voice: Expressive zero-shot singing voice synthesis via midi-driven priors.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference Midi-voice: Expressive zero-shot singing voice synthesis via midi-driven priors

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.099283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.867096Z digest=sha256:79862d80c2fa718c4f10555bc4fb59c7e395d230aca50da46a1e29ffb3c63d56

Observation 2b7caab1-5660-4e00-8121-3c59df2f1e7f · outbound

This paper cites Zero-shot singing voice synthesis from musical score.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference Zero-shot singing voice synthesis from musical score

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.087757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.870646Z digest=sha256:55ce69521f61d6e78283f2a6ab9e939de465488ca558e1eb88370b6b6f195dc3

Observation c1ab64ee-ba1b-4b3d-ab3f-335e3b9d454d · outbound

This paper cites Expressivesinger: Multilingual and multi- style score-based singing voice synthesis with expressive performance control.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference Expressivesinger: Multilingual and multi- style score-based singing voice synthesis with expressive performance control

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.075978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.874394Z digest=sha256:ce95bcef00731c492830e386f0789c43a1c47e8fe01135f20d214fafaae28b0a

Observation cb048025-860e-4a77-a6da-434a7b91fa0a · outbound

This paper cites Music style transfer: A position paper.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference Music style transfer: A position paper

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.064532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.877958Z digest=sha256:ba4f34accf3134f9a4100d8d23214686c89170b47e82513dc9eb05b979ec9123

Observation 9ed4f525-7cb4-4ea3-9eab-157903210a62 · outbound

This paper cites A unified model for zero-shot singing voice con- version and synthesis.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference A unified model for zero-shot singing voice con- version and synthesis

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.051902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.881319Z digest=sha256:8022bbe2cce4191921f2c271554564b11540059c7e0abd82da848b8259acbf2d

Observation 05cbca40-61bf-4eb7-aefc-5ce3701351de · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T15:33:04.885124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:33:04.885124Z digest=sha256:ac917c7f79882f8c5335637797474c33e3512e2cbb3753bbd515317037341ce1

Observation 111cb512-15aa-4359-8f09-02b7551766ca · outbound

This paper cites Generalized end-to-end loss for speaker verifi- cation.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference Generalized end-to-end loss for speaker verifi- cation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.040165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.889517Z digest=sha256:85029b54baf49a2029855846785f2d0d281379c2b1efeb7cd46dc26bd35b94c4

Observation e69c6172-a4bb-47a2-9647-12dc649045ae · outbound

This paper cites Ddsp: Differentiable digital signal processing.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference Ddsp: Differentiable digital signal processing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.028027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.894802Z digest=sha256:e083756190c3c8eaceb08a4a4139ad3d37e4cdf20b45e06b15716bb83f0b269c

Observation 79484e40-908d-4c32-b04c-47cb1807aca1 · outbound

This paper cites wav2vec 2.0: A framework for self- supervised learning of speech representations.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference wav2vec 2.0: A framework for self- supervised learning of speech representations

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.016006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.898524Z digest=sha256:9d6bc2d33bd314aa22dfbbf357f791a7e20d19bd62e4e60db40294d1d11c6ca0

Observation e72ca49b-5ead-41c3-a4fe-818bb6804e01 · outbound

This paper cites Dannenberg.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference Dannenberg

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:33:05.002634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T15:33:04.902133Z digest=sha256:fc8e94d47fe445b7de9de8a58a8b7795bc16d95c703778551b53128f1d12ba27

Observation c007c5c4-a1d1-4b07-821e-08a9aeac419f · outbound

This paper cites LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference LibriTTS-R: A Restored Multi-Speaker Text-to-Speech Corpus

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T15:33:04.905829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:33:04.905829Z digest=sha256:2ca54e6e5034513c77756d703b83206915dba5ba1168f6af18c87757d4b56472

Observation b2ddef1b-fd89-4250-b25c-b5f928ef472f · outbound

This paper cites Denoising Diffusion Implicit Models.

Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference Denoising Diffusion Implicit Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T15:33:04.909705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:33:04.909705Z digest=sha256:dbe528e3b300338f7a18db63509ad78fb35eec509b8545dd7fbb3be1418dbd21

Pith citing papers

No inbound Pith citation observations are available.