Pith. sign in

Paper Citation Record · LEDGER

Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:1712.05884.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1712.05884 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:32:40.409378Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T12:29:51.973055Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e84fefc7-76fb-40d8-9646-0a3adc0e46da · inbound

RUSLAN: Russian Spoken Language Corpus for Speech Synthesis cites this paper.

RUSLAN: Russian Spoken Language Corpus for Speech Synthesis Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-25T15:15:58.483819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-25T15:15:38.672352Z digest=sha256:de00704b237198e73444f0180e68d112ab955ea3c3c30499563961de59b52f3d

Observation fe3e802c-472a-443e-8437-7637d6c83b45 · inbound

Hierarchical Sequence to Sequence Voice Conversion with Limited Data cites this paper.

Hierarchical Sequence to Sequence Voice Conversion with Limited Data Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-24T21:34:58.803771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-24T21:30:16.923017Z digest=sha256:492fd368c28f15bfb03151e97565df016d3529ce83cca4388f1c27344197d4b0

Observation cc08e4ad-b405-4ca7-a066-d8804c3554de · inbound

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck cites this paper.

Unpaired Image-to-Speech Synthesis with Multimodal Information Bottleneck Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-14T12:32:40.409378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:32:40.409378Z digest=sha256:021b7f34904c5a9ba2fb518d6cce26719bffb80ec31431b35327ea97211db2ac

Observation 24e6162f-c0ef-409a-837a-d5076b57dfee · inbound

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model cites this paper.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.890550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.890550Z digest=sha256:1fd04515725f4055ed5c9dade92ddfb3877ee6845e7bf0fc8abbb1cd7d12e687

Observation 6e0926eb-cc97-40b3-9b3d-c95ecd11022c · inbound

PASS: Presentation Automation for Slide Generation and Speech cites this paper.

PASS: Presentation Automation for Slide Generation and Speech Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:03:46.485385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:03:46.485385Z digest=sha256:9da471a946e46912af59c03413b1f28828d5c63c6d2aff24313ed34e97c32465

Observation 3eee1f02-b148-45ec-823e-ce31c42c0ba5 · inbound

Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages cites this paper.

Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:30.947024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:14:30.947024Z digest=sha256:cd3fbbb18ac7acc50c282a5835d36bc19250e1958e7cfd0a5cbb0f2f1d6f0bcc

Observation fece8a9d-292e-4c75-adf3-7a32f0a38a54 · inbound

Mechanisms of Misgeneralization in Physical Sequence Modeling cites this paper.

Mechanisms of Misgeneralization in Physical Sequence Modeling Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

Reference 135

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:44:48.678443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-21T07:44:37.810511Z digest=sha256:b79c25cb288d57ea358ae35275c8f5a7dac0e914abc26228b1f51ecefa39cb0c

Observation 8ee4d21d-4367-4fff-b652-71fa6f57f589 · inbound

The Watermark Shortcut: How Provenance Marking Sabotages Audio Deepfake Detection cites this paper.

The Watermark Shortcut: How Provenance Marking Sabotages Audio Deepfake Detection Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-04T12:29:51.974202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T06:45:08.471913Z digest=sha256:770e3331aca9fb258a0d1e174f24140fc81165bf7fd41eb597d78def5b186e21

Observation 12333eaf-8c9b-4265-8eb5-e5a2924e9724 · inbound

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies cites this paper.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.101601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.101601Z digest=sha256:2917b044100ecbe2a4e63a1020cca4f76e48b908a5484a941ff4671ecbf4b35c