Pith. sign in

Paper Citation Record · LEDGER

Fine-grained robust prosody transfer for single-speaker neural text-to-speech

As of 12 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:1907.02479.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1907.02479 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-25T08:27:16.910352Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-25T08:27:16.910352Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-25T08:30:32.052997Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact8
  • verified fuzzy20
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation db14eace-953b-4c5e-8a29-857535ab6815 · outbound

This paper cites Fine-grained robust prosody transfer for single-speaker neural text-to-speech.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Fine-grained robust prosody transfer for single-speaker neural text-to-speech

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T08:30:32.055801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:2422cef4555100eae1fa2b8d821f9a49adfd0c5844ee164bb4329f5c4e4088b6

Observation e8cb53c8-2f6b-4e98-910f-da6ad4326ee6 · outbound

This paper cites First, a seq2seq acoustic model predicts mel-spectrograms from a sequence of phoneme-level linguistic inputs.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech First, a seq2seq acoustic model predicts mel-spectrograms from a sequence of phoneme-level linguistic inputs

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.825003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:2e7483064dc6a8d21aaccc4efc7bb7e6eef9e072a051210ce90e2f411a0fa86b

Observation 10c8ac42-1939-4764-a3ba-a86fe50656b9 · outbound

This paper cites Then we show the application of V AE for better gener- alization towards unseen speakers.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Then we show the application of V AE for better gener- alization towards unseen speakers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.821104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:fcfd3792651628f9df80238632da51d94161d03962f387a641692900c259cb82

Observation 5f7247ca-381d-49f3-aa99-9692af4f9f54 · outbound

This paper cites Data We conducted experiments on an internal US English dataset of audiobook recordings.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Data We conducted experiments on an internal US English dataset of audiobook recordings

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.828768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:4108726f05c57c66641195b9bf5a2ed0fea80e4854c0af886ce429aa319de173

Observation fd364e43-9682-40eb-acb5-dc9e530d83e5 · outbound

This paper cites an unresolved cited work.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-25T09:06:52.801181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:d86da822bc67785318eacab18e8a2fbdd6210aa0653b8805765656fb1f70d457

Observation 7e5c9f51-beac-4432-b48d-cd60e92fda42 · outbound

This paper cites Char2wav: End-to-end speech syn- thesis.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Char2wav: End-to-end speech syn- thesis

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.751087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:d2c448c8934efb8442cdeafd0f3e1a332b8194feba5b34da72ab93e75a460e18

Observation ba2f932b-15d0-4687-a950-2f75ab143e2d · outbound

This paper cites Natural tts synthesis by conditioning wavenet on mel spectrogram predic- tions.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Natural tts synthesis by conditioning wavenet on mel spectrogram predic- tions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.804848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:33580633ebfad8018427e85149befeb054d763b0f9b548de80b4b835a8280673

Observation e0241bba-fc47-41ad-9741-0525428e9a66 · outbound

This paper cites Neural Speech Synthesis with Transformer Network.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Neural Speech Synthesis with Transformer Network

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:30:32.074602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:89cbce01851545db131f15d2996c51c2376ea8ad85eab3694b77d299bc19f706

Observation a5731762-99ae-4f8b-bb8b-572c2084a752 · outbound

This paper cites In other news: A bi-style text-to-speech model for synthesizing newscaster voice with limited data.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech In other news: A bi-style text-to-speech model for synthesizing newscaster voice with limited data

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.809071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:18a597291a599a7ff352129da09993061525cab096f15424a986a0e762aa6ee2

Observation 45af1267-9b7b-430b-aaa7-c4887636866b · outbound

This paper cites Towards end-to-end prosody transfer for expressive speech synthesis with tacotron.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Towards end-to-end prosody transfer for expressive speech synthesis with tacotron

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.812969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:5283a48774f95aabf9c64384e63dd7c9d57cb67b20c24683bb6177ad3a21695e

Observation 5a2e7b7c-38bc-4504-a639-494fddefaa95 · outbound

This paper cites Deep voice 2: Multi-speaker neural text- to-speech.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Deep voice 2: Multi-speaker neural text- to-speech

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.789830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:ec33f5a4d6a48ff0cbd15a8c4d89f0b150bf24851d4fb5440fcbd66141626a16

Observation 25029453-e096-4e0b-aa56-a45ce9a79918 · outbound

This paper cites Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:30:32.079703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:66dfd983f2f282c4e1bcc3846643ed0e41dc6e0884c9c2d935952e2d7c11ea1a

Observation b51900ce-7714-46a4-911c-88ba23dbb501 · outbound

This paper cites Learning latent representations for style control and transfer in end-to-end speech synthesis.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Learning latent representations for style control and transfer in end-to-end speech synthesis

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:30:32.062090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:ecc7b899485005f4ee6b3536ad68783286b1017c39011e50236985c5a48f2243

Observation 705c5aa8-eb6a-4bcb-bea9-85d744b68617 · outbound

This paper cites Tacotron: Towards end-to-end speech synthesis.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Tacotron: Towards end-to-end speech synthesis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.797960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:54f422049fe005372fb2b69a8a72fe2ace91b60d57ead8106bcdf367ffd39a5b

Observation 6970d838-1b04-4122-9908-22f5bde734e2 · outbound

This paper cites Robust and fine-grained prosody control of end-to-end speech synthesis.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Robust and fine-grained prosody control of end-to-end speech synthesis

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:30:32.050302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:61b09ea11016f8ee64f258a100518afffede66b356f7bbb6834bba4dd44d43bd

Observation 23725094-7cac-4b78-9675-6bb762d0ae28 · outbound

This paper cites Automatic segmentation and labeling of speech.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Automatic segmentation and labeling of speech

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.747394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:9fb24f8f579fd868ae6e187d8cb321a159697b437097d25180c7595f729a0e9f

Observation da69ff2a-1cb1-4359-b740-d7e796c4a056 · outbound

This paper cites Towards achieving robust universal neural vocoding.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Towards achieving robust universal neural vocoding

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:30:32.067762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:1a34d5269656fbbdfafbd3eaff6748a034f39246af0f5d0724196001e9e16ce8

Observation 3c811f8e-74bb-483d-884c-4c6a9045634a · outbound

This paper cites Effect of data reduction on sequence-to-sequence neural tts.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Effect of data reduction on sequence-to-sequence neural tts

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.817237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:2637678d52788d84846d670775ced0c02cf9ed6d5fbd4604654357d6ba4d8646

Observation 90ff4946-860a-485c-94b2-d6c5a987be36 · outbound

This paper cites Auto-Encoding Variational Bayes.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Auto-Encoding Variational Bayes

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:30:32.045044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:fc1bcb57b1137782406f2d1c27d9c12ed9bf76b46869b8842c61fb8c4f374a09

Observation 448aca12-f537-4413-b74c-917890d842e7 · outbound

This paper cites Expressive Speech Synthesis via Modeling Expressions with Variational Autoencoder.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Expressive Speech Synthesis via Modeling Expressions with Variational Autoencoder

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:30:32.088042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:e2514c2db7a61b5581b6eda45277d6a6ac7736644bc4f21543b64f4faa8ac3e8

Observation d914be73-ccf9-49ea-b3a1-1dc88698e5dd · outbound

This paper cites On adaptive control processes.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech On adaptive control processes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.777639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:a3fca2e34424e900f3532fe12a2d905674626433ee7b6bba08d86924e821b1d7

Observation 287c9344-cc8c-4d79-8b27-d7c144ad5ce2 · outbound

This paper cites Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced clas- sification frontend.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced clas- sification frontend

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.769824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:1e12da47b615e854e793b17aaad0adaf68737d3fd01ed6bc100c5b350aa89a17

Observation 58bd9de0-0024-4712-ba7b-0cb22f407564 · outbound

This paper cites Joint robust voicing detection and pitch estimation based on residual harmonics.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Joint robust voicing detection and pitch estimation based on residual harmonics

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.773957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:b70d3e73ec15203e2802745f55ae94f4e376714c43b75f3684b75a0cb8e4e359

Observation 25749930-db4d-42df-8142-177264016ef2 · outbound

This paper cites 1534-1,method for the subjective assessment of in- termediate quality levels of coding systems (mushra).

Fine-grained robust prosody transfer for single-speaker neural text-to-speech 1534-1,method for the subjective assessment of in- termediate quality levels of coding systems (mushra)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.781845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:fdae080ece35d268b2e98e7c41b60cef04579dc3fb4a40b966ecc91a8483c65f

Observation bd83952d-8057-4310-bfc8-a2c4a2734bb0 · outbound

This paper cites Statistical analysis of the blizzard challenge 2007 listening test results.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Statistical analysis of the blizzard challenge 2007 listening test results

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.785929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:481fce4b7ff080d1c474750b25918af8a5fd6d7bd7be6f92c9e7da3fd687d390

Observation e15b316d-087b-4538-9b01-133821962b55 · outbound

This paper cites Phonetic poste- riorgrams for many-to-one voice conversion without parallel data training.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Phonetic poste- riorgrams for many-to-one voice conversion without parallel data training

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.794416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:297110dc691006c53ed4235114822ad106cfa44b1c7b03ab108fb6fc82bdd81f

Observation f3beba7d-a8a7-44c9-9bde-3ff744f2b647 · outbound

This paper cites V oice con- version across arbitrary speakers based on a single target-speaker utterance.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech V oice con- version across arbitrary speakers based on a single target-speaker utterance

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.765523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:10829b943a9fdd3050a343f6737f85e7517ea0bd69b3c3251080c047d4a7c6d0

Observation 51ff81b0-f8f9-42b3-8f3c-dd60da89375a · outbound

This paper cites Deep Speech: Scaling up end-to-end speech recognition.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Deep Speech: Scaling up end-to-end speech recognition

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:30:32.083727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:4ca2937cb3e1dfce04d2785b189a2b2d9b5fc86980715b137ee5fa9bbbf213b7

Observation 57f7692d-7d9c-41fc-8d17-efb6cea94f21 · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Lib- rispeech: an asr corpus based on public domain audio books

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.758651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:ae64b7e97a115bd4c8632f50fb1ef5eaf2174287011b6bac42edacf124583238

Observation 820d2561-c44d-467e-b4fe-4bca19f0679f · outbound

This paper cites Common voice.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Common voice

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.754685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:e68aaae3ef39e4ccf5ef098de572c643279fc93bcc289121cd5df0a24775ce1b

Pith citing papers

Observation db14eace-953b-4c5e-8a29-857535ab6815 · inbound

Fine-grained robust prosody transfer for single-speaker neural text-to-speech cites this paper.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Fine-grained robust prosody transfer for single-speaker neural text-to-speech

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T08:30:32.055801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:2422cef4555100eae1fa2b8d821f9a49adfd0c5844ee164bb4329f5c4e4088b6