Pith. sign in

Paper Citation Record · LEDGER

A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction

As of 12 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2412.08312.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08312 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:01:14.189712Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact3
  • verified fuzzy2
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation affa0bbf-aa0c-41c6-9d4c-edc366dbaa46 · outbound

This paper cites V oice Conversion Using Speech-to-Speech Neuro-Style Transfer,.

A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction V oice Conversion Using Speech-to-Speech Neuro-Style Transfer,

Reference 1

Resolution
verified exact
doi, observed 2026-08-11T18:01:14.239345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:01:14.150058Z digest=sha256:554a39732330567dda25c85de88defcad8cb6ec84745b40a70e8592c6afaef98

Observation 65d994bd-5d78-46e2-81c0-712836bda15b · outbound

This paper cites Hidden Markov Model based Speech Synthesis: A Review,.

A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction Hidden Markov Model based Speech Synthesis: A Review,

Reference 2

Resolution
verified exact
doi, observed 2026-08-11T18:01:14.228230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:01:14.155113Z digest=sha256:429af24faecb07dff4897618d959c9431912788a5bb4c294a4a10b07652cd769

Observation 7c77385d-2cfa-4a3e-87ab-510a47805779 · outbound

This paper cites Phoneme independent HMM voice conversion,.

A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction Phoneme independent HMM voice conversion,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T18:01:14.158669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:01:14.158669Z digest=sha256:83321164e712b22388b7c2f6c1023e4cf368ef4d5281e1b8d8d99433c2d388b3

Observation 90bfa4a8-776e-4ceb-8586-f4d93b773065 · outbound

This paper cites Speaker Conditional WaveRNN: Towards Universal Neural Vocoder for Unseen Speaker and Recording Conditions.

A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction Speaker Conditional WaveRNN: Towards Universal Neural Vocoder for Unseen Speaker and Recording Conditions

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-11T18:01:14.554582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:01:14.169233Z digest=sha256:c34456f88907c13d275a0c921b879ba6a8dfb0c187d9ed3e0236375def7b8f51

Observation 1e29502a-4ac9-4648-9efb-8095c9b5a254 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:01:14.664464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:01:14.172936Z digest=sha256:43aef2e31ae051bceef2cf133bde768dc23a37ba8e0e2153621294890aa5c207

Observation 81cec743-1e60-4d39-9f40-b9812f65c084 · outbound

This paper cites Pitchnet: Unsupervised Singing V oice Conversion with Pitch Adversarial Network,.

A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction Pitchnet: Unsupervised Singing V oice Conversion with Pitch Adversarial Network,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:01:14.176542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:01:14.176542Z digest=sha256:750534f36f8e72551bf628a761c34e89958445c8406ecc2bc5411a594340c11e

Observation 1c988329-89a3-4978-b8d4-eddebdb09724 · outbound

This paper cites The Singing V oice Conversion Challenge 2023,.

A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction The Singing V oice Conversion Challenge 2023,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T18:01:14.180207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:01:14.180207Z digest=sha256:8a42576a0eeb14aef14ebebfef12e99da0f61c745cf0595b7dfdde607db862f8

Observation 6f8c64be-2ab8-4f7c-adc2-3729cc37203e · outbound

This paper cites Accent modification for speech recognition of non-native speakers using neural style transfer,.

A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction Accent modification for speech recognition of non-native speakers using neural style transfer,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T18:01:14.183333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:01:14.183333Z digest=sha256:c1fea390cf17d71a3b8d3818ecd71654bdb16316be92173b5c03f922d3d2748a

Observation c53006fd-4790-4d99-a374-4bf61e3464a2 · outbound

This paper cites Parallel voice conversion with limited training data using stochastic variational deep kernel learning,.

A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction Parallel voice conversion with limited training data using stochastic variational deep kernel learning,

Reference 10

Resolution
metadata mismatch
raw_fallback, observed 2026-08-11T18:01:14.321204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:01:14.186523Z digest=sha256:b6de7f903aee97f8718f61b1f1435e5517463aeb3e5f3d9ce12d65cdd45a1002

Observation 87d2e52d-d352-4f62-86f8-998a3a8a111a · outbound

This paper cites Speech Accent Archive,.

A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction Speech Accent Archive,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:01:14.653144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T18:01:14.189712Z digest=sha256:805967bb238a9a6aa71465be7b0dcb422b88d219adc708dcb2939fbaa841df94

Observation 53f5181f-6cc7-4fd9-b5e0-d53044501f8d · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

A Unified Model For Voice and Accent Conversion In Speech and Singing using Self-Supervised Learning and Feature Extraction WaveNet: A Generative Model for Raw Audio

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-11T18:01:14.165761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:01:14.165761Z digest=sha256:e0d27771e78d7eef8d373b9c00811e18f8f6b34c4a03e8b9c6f3576c45de1306

Pith citing papers

No inbound Pith citation observations are available.