Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:16:32.475093Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2506.02997.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:16:32.475093Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
18 of 18 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 98ee0846-7ee5-4696-a363-c8518914cb4d · outbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Prompttts: Controllable text-to-speech with text descriptions,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43e8d4ab-718e-47cb-bf3e-fabe5fd92d5d · outbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation PromptTTS 2: Describing and Generating Voices with Text Prompt
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01f03004-b063-4a86-81c5-1c4eb5d14d2c · outbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Textrolspeech: A text style control speech corpus with codec language text-to-speech models,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d2bd77f-ec90-4f34-8bdb-9bbeb17e9ad7 · outbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4274f1ba-67a4-4902-ae9d-8cb9510965e3 · outbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 36e3e659-0c81-48b7-8ddc-5fbf05c4dd55 · outbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Prosody-tts: Improving prosody with masked autoencoder and conditional diffusion model for expressive text-to-speech,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7fc73b80-3143-46a8-8b40-71a6b50a205b · outbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Uniaudio: Towards universal audio generation with large language models,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 71595cf2-7d84-4261-bb37-a801c98d8cf4 · outbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Classifier-free diffusion guidance,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c4cdf49-992a-4d22-a6ff-09981ad38bb5 · outbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 827789aa-6c87-4b25-b9ce-1b1bc304e76f · outbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Librispeech: an asr corpus based on public domain audio books,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb796082-112f-4c2d-b715-ee251fdd9099 · outbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6af88a1a-5f9a-498a-b07e-3a8ffabb8bc3 · outbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Dailytalk: Spoken dialogue dataset for conversational text-to-speech,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f3069eef-0cf1-4c51-be97-96203ce4197e · outbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c89fd632-0615-417d-824b-23b60a90d21b · outbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Robust Speech Recognition via Large-Scale Weak Supervision
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d45114f-5714-4f3e-a082-87767117d8d6 · outbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation High Fidelity Neural Audio Compression
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27a2671c-87a4-4c6d-a93e-23daa4105847 · outbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 975c135b-5556-4155-bad1-0e7301bc06e9 · outbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4d329d4-8d54-4a05-b4da-311e7867885a · outbound
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Wespeaker: A research and production oriented speaker embedding learning toolkit,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.