Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T22:51:50.050472Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2412.03074.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T22:51:50.050472Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 28aeab77-1755-46f0-a88a-d109bc6cc833 · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model A Brief Overview of Unsupervised Neural Speech Representation Learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24e6162f-c0ef-409a-837a-d5076b57dfee · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff9c97c0-aaa2-497e-b7b9-dd602b6bb6bc · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2596427f-2627-40a1-8ddc-01f464ec602d · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6b041915-aab1-493c-951d-0ce34066dfd5 · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Natural TTS synthesis by conditioning Wavenet on mel-spectrogram predictions,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4914c130-3b3c-4289-a90b-8357c6324bc6 · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model On generative spoken language modeling from raw audio,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9f2b912-7f0a-48b6-8c94-7df3355f3f90 · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Representation Learning with Contrastive Predictive Coding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fd38301-1d33-4570-8c13-dad36764ceb5 · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Wav2vec 2.0: A framework for self-supervised learning of speech representations,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a807a468-54ee-468e-a89b-502e821a3a4e · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model HuBERT: Self- supervised speech representation learning by masked prediction of hidden units,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 424690e8-0ec4-49c8-bbfc-5cd567793e8c · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Attention is all you need,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 62613339-6e88-422d-b6f7-bb67e2a1afbc · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model The Zero Resource Speech Challenge 2020: Discovering Discrete Subword and Word Units,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 426bac3e-2054-43a0-9b94-25b6876fb6bf · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model The zero resource speech challenge 2021: Spoken language mod- elling,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9e2a6a93-268c-4b36-aaeb-7d701ee17560 · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Self- supervised language learning from raw audio: Lessons from the zero resource speech challenge,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation afa51aa1-87a9-461f-9db8-b84826267a1a · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model A comparison of discrete and soft speech units for improved voice conversion,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b10deed6-9dcb-4f28-95ce-5f99ed039354 · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Textless direct speech-to- speech translation with discrete speech representation,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation af9ee82f-fc6c-41e9-8eba-ca1a7d11538d · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation be39764a-d380-48bb-bc88-d67353c8bfb5 · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model LibriSpeech: An ASR corpus based on public domain audio books,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3b8d78e1-02aa-4d7d-9084-518504527449 · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Ito and L
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 26165c73-cf25-42f4-95ea-76ade09ae3eb · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9af1f3d6-57ea-4a1c-979a-6a410715f544 · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aba34485-f38b-4812-a5ff-9bf6ab45dbce · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model JVS corpus: free Japanese multi-speaker voice corpus
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 084663e5-cfc0-46c9-b2b8-d3d1fff071c9 · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Audio- book speech synthesis conditioned by cross-sentence context-aware word embeddings,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 904682ab-e720-4882-9125-c6a52e1496ea · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model J-MAC: Japanese multi-speaker audiobook corpus for speech synthesis,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 27ce4de1-9389-43c3-ae3a-80f37022a628 · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Radford, K
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b2727b2-2a2b-486f-a526-48cf884fe8a8 · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Investi- gation of robustness of hubert features from different layers to domain, accent and language variations,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation af4827dd-6f50-4f16-99e2-0a9732380725 · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model ContentVec: An improved self-supervised speech representation by dis- entangling speakers,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c61e9b87-1dbc-4368-81d9-e557c5c21a8c · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model JSSS: free Japanese speech corpus for summarization and simplification
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3828493d-eb62-45b8-a824-91f186f82db4 · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Universal phone recognition with a multilingual allophone system,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7aa3af17-3970-46e3-8975-52bbd7399cf4 · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Speech quality assessment with W ARP-Q: From sim- ilarity to subsequence dynamic time warp cost,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a362ed75-5413-4e62-9c3f-f8c4e3014f7f · outbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model Unresolved cited work
Reference 3042
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
No inbound Pith citation observations are available.