Pith. sign in

Paper Citation Record · LEDGER

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model

As of 14 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2412.03074.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.03074 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:51:50.050472Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact1
  • verified fuzzy15
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 28aeab77-1755-46f0-a88a-d109bc6cc833 · outbound

This paper cites A Brief Overview of Unsupervised Neural Speech Representation Learning.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model A Brief Overview of Unsupervised Neural Speech Representation Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.884332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.884332Z digest=sha256:8c97de2587d1c3c850b89b1bb521c7ff63865c613a77f3efbd770119c8ad9cf8

Observation 24e6162f-c0ef-409a-837a-d5076b57dfee · outbound

This paper cites Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.890550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.890550Z digest=sha256:2ba08f021c71e254de6fdcbee3de2fb282f9dff9a8e8d70afe565eb5eb17ff60

Observation ff9c97c0-aaa2-497e-b7b9-dd602b6bb6bc · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.895990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.895990Z digest=sha256:dc0b0833c278e25bcd7bc2bb6582e22f6007dae27236ad7b139d09480c80aa19

Observation 2596427f-2627-40a1-8ddc-01f464ec602d · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.610616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:51:49.901188Z digest=sha256:ee54f9aa1a22c79adbb5c41535c6dd6950b9599097ac32222219256ba51ac6d0

Observation 6b041915-aab1-493c-951d-0ce34066dfd5 · outbound

This paper cites Natural TTS synthesis by conditioning Wavenet on mel-spectrogram predictions,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Natural TTS synthesis by conditioning Wavenet on mel-spectrogram predictions,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.906849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.906849Z digest=sha256:38fe4a6f39ba34e57bd9233ca230b6ada4b96da6470060f38b0422cd4dcd1dd9

Observation 4914c130-3b3c-4289-a90b-8357c6324bc6 · outbound

This paper cites On generative spoken language modeling from raw audio,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model On generative spoken language modeling from raw audio,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.912368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.912368Z digest=sha256:f131f66d28aa4241f96f9fc186b93f18c1796b1a8675eefdb645bb72183a92fe

Observation c9f2b912-7f0a-48b6-8c94-7df3355f3f90 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Representation Learning with Contrastive Predictive Coding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.918052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.918052Z digest=sha256:a7a267857673b3cc9fe8efdfff2d24f5e3e4871d2c19c721c1d6984ecc49d75c

Observation 7fd38301-1d33-4570-8c13-dad36764ceb5 · outbound

This paper cites Wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.926589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.926589Z digest=sha256:5c4f98e194c8658c12bcc1c1389592af884809257d6ed2da4c6fb1bbf0951fad

Observation a807a468-54ee-468e-a89b-502e821a3a4e · outbound

This paper cites HuBERT: Self- supervised speech representation learning by masked prediction of hidden units,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model HuBERT: Self- supervised speech representation learning by masked prediction of hidden units,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.932289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.932289Z digest=sha256:2645f9bcd9efefd6c688ba941bfd2249a3147f5c841adcdb530c2f7e973238cb

Observation 424690e8-0ec4-49c8-bbfc-5cd567793e8c · outbound

This paper cites Attention is all you need,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Attention is all you need,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.532944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:51:49.937783Z digest=sha256:5edb17e1169a1df4453e2ab070e7cc12af2cd5721a8851eda875721144dc1de4

Observation 62613339-6e88-422d-b6f7-bb67e2a1afbc · outbound

This paper cites The Zero Resource Speech Challenge 2020: Discovering Discrete Subword and Word Units,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model The Zero Resource Speech Challenge 2020: Discovering Discrete Subword and Word Units,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.513108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:51:49.943099Z digest=sha256:717b2cb5085cec41b5e3e5634aee735376445c1d716db2f92cfd8acc42d120c4

Observation 426bac3e-2054-43a0-9b94-25b6876fb6bf · outbound

This paper cites The zero resource speech challenge 2021: Spoken language mod- elling,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model The zero resource speech challenge 2021: Spoken language mod- elling,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.494312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:51:49.948890Z digest=sha256:c777f1ee4803cad5b608f1f689448a50632bb46be88e0b97efe2f3b96a8d803d

Observation 9e2a6a93-268c-4b36-aaeb-7d701ee17560 · outbound

This paper cites Self- supervised language learning from raw audio: Lessons from the zero resource speech challenge,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Self- supervised language learning from raw audio: Lessons from the zero resource speech challenge,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.476981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:51:49.954166Z digest=sha256:24e24b1ad764aebba732edaf55b0289009ff9f71b3dadf33cd0783dc101c0958

Observation afa51aa1-87a9-461f-9db8-b84826267a1a · outbound

This paper cites A comparison of discrete and soft speech units for improved voice conversion,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model A comparison of discrete and soft speech units for improved voice conversion,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.457515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:51:49.958904Z digest=sha256:1c2fb64cae1f54e1e57794414ea38b875c787b684d5a44238472a7ab6bdd17b1

Observation b10deed6-9dcb-4f28-95ce-5f99ed039354 · outbound

This paper cites Textless direct speech-to- speech translation with discrete speech representation,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Textless direct speech-to- speech translation with discrete speech representation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.436813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:51:49.963908Z digest=sha256:37db0b77eb4e4a7ddfa412486e5db2680b411faea31d20db0b69a3fd64baaa2c

Observation af9ee82f-fc6c-41e9-8eba-ca1a7d11538d · outbound

This paper cites UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.414008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:51:49.969832Z digest=sha256:5ae55208717e6c66692a558db868c844bf22c824f559be37c7f4a72ddb4e6752

Observation be39764a-d380-48bb-bc88-d67353c8bfb5 · outbound

This paper cites LibriSpeech: An ASR corpus based on public domain audio books,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model LibriSpeech: An ASR corpus based on public domain audio books,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.394880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:51:49.982938Z digest=sha256:48df77d9bef95c2e232f220187ddcc914c40678195f7c4fa1d416a16472015f5

Observation 3b8d78e1-02aa-4d7d-9084-518504527449 · outbound

This paper cites Ito and L.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Ito and L

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.373300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:51:49.987883Z digest=sha256:5a7cde60c9898b644b8f4854f7be3dfb6f7944625a1602042babe7db4172b9d4

Observation 26165c73-cf25-42f4-95ea-76ade09ae3eb · outbound

This paper cites an unresolved cited work.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-11T22:51:50.354281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:51:49.992872Z digest=sha256:fc7fed936e16afc30926340acd59533e438d4458ccf81aea6c854ed1658f13e7

Observation 9af1f3d6-57ea-4a1c-979a-6a410715f544 · outbound

This paper cites JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:49.998068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:49.998068Z digest=sha256:59473a40880617b26b3803b6c537e1e3d01d759f1032f45b0c5090c232c9de5c

Observation aba34485-f38b-4812-a5ff-9bf6ab45dbce · outbound

This paper cites JVS corpus: free Japanese multi-speaker voice corpus.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model JVS corpus: free Japanese multi-speaker voice corpus

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:50.005341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:50.005341Z digest=sha256:e5cbb504b152d277283c5703bc586129cae2a179a7f67024bec7fd2884b8ed52

Observation 084663e5-cfc0-46c9-b2b8-d3d1fff071c9 · outbound

This paper cites Audio- book speech synthesis conditioned by cross-sentence context-aware word embeddings,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Audio- book speech synthesis conditioned by cross-sentence context-aware word embeddings,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:50.011574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:50.011574Z digest=sha256:b82f5dae78acbfb0e094e94991590b50ac17822bf635169813ba3b8348f403e1

Observation 904682ab-e720-4882-9125-c6a52e1496ea · outbound

This paper cites J-MAC: Japanese multi-speaker audiobook corpus for speech synthesis,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model J-MAC: Japanese multi-speaker audiobook corpus for speech synthesis,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.323938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:51:50.016467Z digest=sha256:5a9d001238106a5266d1ad202eea68f0e5aa8c76d4732795dc5f4fa529b7bb3f

Observation 27ce4de1-9389-43c3-ae3a-80f37022a628 · outbound

This paper cites Radford, K.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Radford, K

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:50.022731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:50.022731Z digest=sha256:f507c71ac231e6a88890b5bc7e366b75219900f35139bb6a90c7eaed9a8d9419

Observation 7b2727b2-2a2b-486f-a526-48cf884fe8a8 · outbound

This paper cites Investi- gation of robustness of hubert features from different layers to domain, accent and language variations,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Investi- gation of robustness of hubert features from different layers to domain, accent and language variations,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.294196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:51:50.027954Z digest=sha256:934ca77e0cf9d049adb1777880cc53d944ea88abfb685252e6098d1f3d5edb23

Observation af4827dd-6f50-4f16-99e2-0a9732380725 · outbound

This paper cites ContentVec: An improved self-supervised speech representation by dis- entangling speakers,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model ContentVec: An improved self-supervised speech representation by dis- entangling speakers,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.277654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:51:50.033061Z digest=sha256:42fa84bdd9dde3d82119b53132049662069d91c6b80aee38a2b6f74fa70122de

Observation c61e9b87-1dbc-4368-81d9-e557c5c21a8c · outbound

This paper cites JSSS: free Japanese speech corpus for summarization and simplification.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model JSSS: free Japanese speech corpus for summarization and simplification

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T22:51:50.039933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:51:50.039933Z digest=sha256:d855b005e529a0c26a3745b298282270b2d7d27f0cb341723c649411e07a3bc2

Observation 3828493d-eb62-45b8-a824-91f186f82db4 · outbound

This paper cites Universal phone recognition with a multilingual allophone system,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Universal phone recognition with a multilingual allophone system,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.258903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:51:50.045137Z digest=sha256:58b6c03f610ef7d9027dc585445707411b38b160319f37a621067dbce13c1ca7

Observation 7aa3af17-3970-46e3-8975-52bbd7399cf4 · outbound

This paper cites Speech quality assessment with W ARP-Q: From sim- ilarity to subsequence dynamic time warp cost,.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Speech quality assessment with W ARP-Q: From sim- ilarity to subsequence dynamic time warp cost,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T22:51:50.240815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:51:50.050472Z digest=sha256:c13878937b4a95a3999d8764ad9020ce3efa187dae18894ececa7f873dc4298b

Observation a362ed75-5413-4e62-9c3f-f8c4e3014f7f · outbound

This paper cites an unresolved cited work.

Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Unresolved cited work

Reference 3042

Resolution
verified exact
doi, observed 2026-08-11T22:51:50.091253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T22:51:49.977002Z digest=sha256:4a7c501c367e76a8a04089a53e21888f3e9e6cff97f63f62d932c6e282155275

Pith citing papers

No inbound Pith citation observations are available.