Pith. sign in

Paper Citation Record · LEDGER

Fine-grained robust prosody transfer for single-speaker neural text-to-speech

As of 4 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:1907.02479.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1907.02479 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-25T08:27:16.910352Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-25T08:27:16.910352Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-25T08:30:32.052997Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact8
  • verified fuzzy20
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation db14eace-953b-4c5e-8a29-857535ab6815 · outbound

This paper cites Fine-grained robust prosody transfer for single-speaker neural text-to-speech.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Fine-grained robust prosody transfer for single-speaker neural text-to-speech

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T08:30:32.055801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:a5d1a003435e926c0288dc2fd2a90bd83d143b2f495d027b70a1cacd58e6c6d9

Observation e8cb53c8-2f6b-4e98-910f-da6ad4326ee6 · outbound

This paper cites First, a seq2seq acoustic model predicts mel-spectrograms from a sequence of phoneme-level linguistic inputs.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech First, a seq2seq acoustic model predicts mel-spectrograms from a sequence of phoneme-level linguistic inputs

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.825003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:e7f94653cff982a32de2bd768abff0e4508861fff16ea2afb93aca11ddbaae63

Observation 10c8ac42-1939-4764-a3ba-a86fe50656b9 · outbound

This paper cites Then we show the application of V AE for better gener- alization towards unseen speakers.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Then we show the application of V AE for better gener- alization towards unseen speakers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.821104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:fbefb96cf66d0289290e1ba8c435c32fd55e29c75aecc3b0cb755051fd69d098

Observation 5f7247ca-381d-49f3-aa99-9692af4f9f54 · outbound

This paper cites Data We conducted experiments on an internal US English dataset of audiobook recordings.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Data We conducted experiments on an internal US English dataset of audiobook recordings

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.828768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:5fad5da76ddd767be848aa5d7fd5fbfd49d0e71070cc1bb7ae58f99ec13a0468

Observation fd364e43-9682-40eb-acb5-dc9e530d83e5 · outbound

This paper cites an unresolved cited work.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-25T09:06:52.801181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:fadbd8a9d4bbce1433bef3256fc45c5c61cc14300477243e75b0356a5d8a6cda

Observation 7e5c9f51-beac-4432-b48d-cd60e92fda42 · outbound

This paper cites Char2wav: End-to-end speech syn- thesis.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Char2wav: End-to-end speech syn- thesis

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.751087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:717a3241119aa932ea540dc17278197a9954491ab6216cef22915ad64621aa87

Observation ba2f932b-15d0-4687-a950-2f75ab143e2d · outbound

This paper cites Natural tts synthesis by conditioning wavenet on mel spectrogram predic- tions.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Natural tts synthesis by conditioning wavenet on mel spectrogram predic- tions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.804848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:640bbdbfd438d6b57012ce3f347e00b145ef1544fd1a1db02c6dfa6cfb3411b3

Observation e0241bba-fc47-41ad-9741-0525428e9a66 · outbound

This paper cites Neural Speech Synthesis with Transformer Network.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Neural Speech Synthesis with Transformer Network

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:30:32.074602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:5d7365d80692fdf2546413c5c4c0fdf89c0f07df4c647824d4abd71ec13a7e38

Observation a5731762-99ae-4f8b-bb8b-572c2084a752 · outbound

This paper cites In other news: A bi-style text-to-speech model for synthesizing newscaster voice with limited data.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech In other news: A bi-style text-to-speech model for synthesizing newscaster voice with limited data

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.809071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:ed7c12271b0d5a6816ccd4ccd4b03f50ddd893da1ae1334efb419712b090736f

Observation 45af1267-9b7b-430b-aaa7-c4887636866b · outbound

This paper cites Towards end-to-end prosody transfer for expressive speech synthesis with tacotron.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Towards end-to-end prosody transfer for expressive speech synthesis with tacotron

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.812969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:e4f6ea17b14eac108ecf792f6d3a3e188c2e42d7bb153645468276ad4c7c7f08

Observation 5a2e7b7c-38bc-4504-a639-494fddefaa95 · outbound

This paper cites Deep voice 2: Multi-speaker neural text- to-speech.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Deep voice 2: Multi-speaker neural text- to-speech

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.789830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:d42f4c7ae985cbd7e5ea7e9e0959c2a5afb853421617d5bfddae648cf61be91f

Observation 25029453-e096-4e0b-aa56-a45ce9a79918 · outbound

This paper cites Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:30:32.079703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:f62d5d21f584a25950762b3213c2ffabe8d0c0d4a32332f4f1ae13fa6a84b937

Observation b51900ce-7714-46a4-911c-88ba23dbb501 · outbound

This paper cites Learning latent representations for style control and transfer in end-to-end speech synthesis.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Learning latent representations for style control and transfer in end-to-end speech synthesis

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:30:32.062090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:d7b97024ef6e4e1beb12a4c3d690e2550a7df10801cf2143f2e3fd0ae80a90af

Observation 705c5aa8-eb6a-4bcb-bea9-85d744b68617 · outbound

This paper cites Tacotron: Towards end-to-end speech synthesis.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Tacotron: Towards end-to-end speech synthesis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.797960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:fe3781cea161114d6c2c5954c7512dfd56ccc0e13df92728d7e3befdcff9b7b8

Observation 6970d838-1b04-4122-9908-22f5bde734e2 · outbound

This paper cites Robust and fine-grained prosody control of end-to-end speech synthesis.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Robust and fine-grained prosody control of end-to-end speech synthesis

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:30:32.050302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:6780258b8253b82602d4bfe63beb2e41eb87c517f8e4b1cf32c9d76528c73c6c

Observation 23725094-7cac-4b78-9675-6bb762d0ae28 · outbound

This paper cites Automatic segmentation and labeling of speech.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Automatic segmentation and labeling of speech

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.747394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:8fb02a54a897f08125e5db12c8fc058be1623ac9b2daa306edd72139af54c6dd

Observation da69ff2a-1cb1-4359-b740-d7e796c4a056 · outbound

This paper cites Towards achieving robust universal neural vocoding.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Towards achieving robust universal neural vocoding

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:30:32.067762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:5336659df0b80f1e32548d745d59d7aa26c4d89652d803b904dcb1cf27ca5f4a

Observation 3c811f8e-74bb-483d-884c-4c6a9045634a · outbound

This paper cites Effect of data reduction on sequence-to-sequence neural tts.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Effect of data reduction on sequence-to-sequence neural tts

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.817237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:5d7a8ed5584b033c2dd0da9407bf4a76c2244016245d74bd29d42ec211200fde

Observation 90ff4946-860a-485c-94b2-d6c5a987be36 · outbound

This paper cites Auto-Encoding Variational Bayes.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Auto-Encoding Variational Bayes

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:30:32.045044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:fcf04aaac0356ce64efe9bdfd8bd6abfd20fb83d0e101f1f573a5a0aee58b7ce

Observation 448aca12-f537-4413-b74c-917890d842e7 · outbound

This paper cites Expressive Speech Synthesis via Modeling Expressions with Variational Autoencoder.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Expressive Speech Synthesis via Modeling Expressions with Variational Autoencoder

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:30:32.088042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:83d9780ebe0d88e4a1a77edcf5197396542adb75c33a7243238fe8e45f127515

Observation d914be73-ccf9-49ea-b3a1-1dc88698e5dd · outbound

This paper cites On adaptive control processes.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech On adaptive control processes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.777639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:40cccebc26430dda6542e0faf4e3c440923b929802801ebb4912ca6ae869d95a

Observation 287c9344-cc8c-4d79-8b27-d7c144ad5ce2 · outbound

This paper cites Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced clas- sification frontend.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced clas- sification frontend

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.769824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:af1f28b9a15ebbb7542d1aeb9fb2f2c00069ceff5df5299d0e0831e14dc29c71

Observation 58bd9de0-0024-4712-ba7b-0cb22f407564 · outbound

This paper cites Joint robust voicing detection and pitch estimation based on residual harmonics.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Joint robust voicing detection and pitch estimation based on residual harmonics

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.773957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:824858da83474aea20343b6ddcdad4f0add6627cc665260cf42678127f0f1e1d

Observation 25749930-db4d-42df-8142-177264016ef2 · outbound

This paper cites 1534-1,method for the subjective assessment of in- termediate quality levels of coding systems (mushra).

Fine-grained robust prosody transfer for single-speaker neural text-to-speech 1534-1,method for the subjective assessment of in- termediate quality levels of coding systems (mushra)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.781845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:1760dc75be742b4e63602a8b592848ae609e20650c47526f24f99181674b7b62

Observation bd83952d-8057-4310-bfc8-a2c4a2734bb0 · outbound

This paper cites Statistical analysis of the blizzard challenge 2007 listening test results.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Statistical analysis of the blizzard challenge 2007 listening test results

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.785929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:3dfbbf7643374a334270fde7b898dd009faaddc4724598d0c94cc4c0dd11f0e4

Observation e15b316d-087b-4538-9b01-133821962b55 · outbound

This paper cites Phonetic poste- riorgrams for many-to-one voice conversion without parallel data training.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Phonetic poste- riorgrams for many-to-one voice conversion without parallel data training

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.794416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:64eb28f1ed6fef24262a883bb63b287bf54a6d5723508032808056ec8ba404a3

Observation f3beba7d-a8a7-44c9-9bde-3ff744f2b647 · outbound

This paper cites V oice con- version across arbitrary speakers based on a single target-speaker utterance.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech V oice con- version across arbitrary speakers based on a single target-speaker utterance

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.765523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:2be3cd6ef9a0c06adaca8f613c141c7dfc9fe7524b376f21eb8c24e8b1a8d9c5

Observation 51ff81b0-f8f9-42b3-8f3c-dd60da89375a · outbound

This paper cites Deep Speech: Scaling up end-to-end speech recognition.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Deep Speech: Scaling up end-to-end speech recognition

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:30:32.083727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:2576642b13bd2e17b6d535ca07a11ec1dea0666e4a3a10dd39e7f32c0dfe0fb3

Observation 57f7692d-7d9c-41fc-8d17-efb6cea94f21 · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Lib- rispeech: an asr corpus based on public domain audio books

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.758651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:fe9d8a8fbec090085d56d2f89f328555dce9df0754fb79b98c82480d59a69c83

Observation 820d2561-c44d-467e-b4fe-4bca19f0679f · outbound

This paper cites Common voice.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Common voice

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T09:06:52.754685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:379e7e5852d7b3b1f5b5beda91ce035cbad8641e68e21b87b1fc6556e8a8f999

Pith citing papers

Observation db14eace-953b-4c5e-8a29-857535ab6815 · inbound

Fine-grained robust prosody transfer for single-speaker neural text-to-speech cites this paper.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Fine-grained robust prosody transfer for single-speaker neural text-to-speech

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T08:30:32.055801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:a5d1a003435e926c0288dc2fd2a90bd83d143b2f495d027b70a1cacd58e6c6d9