Pith. sign in

Paper Citation Record · LEDGER

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation

As of 8 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2506.02997.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02997 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:16:32.475093Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98ee0846-7ee5-4696-a363-c8518914cb4d · outbound

This paper cites Prompttts: Controllable text-to-speech with text descriptions,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Prompttts: Controllable text-to-speech with text descriptions,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:34.432358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:16:31.033516Z digest=sha256:f49cecee5896d7050c9b2f23cb6ee57b0e5313792cdc5a7528c92c3c8c88e6e5

Observation 43e8d4ab-718e-47cb-bf3e-fabe5fd92d5d · outbound

This paper cites PromptTTS 2: Describing and Generating Voices with Text Prompt.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation PromptTTS 2: Describing and Generating Voices with Text Prompt

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.114971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.114971Z digest=sha256:72048a46af74303e54f44b8e9118509f63415082a4c06117ecfb3d5fad76ea8a

Observation 01f03004-b063-4a86-81c5-1c4eb5d14d2c · outbound

This paper cites Textrolspeech: A text style control speech corpus with codec language text-to-speech models,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Textrolspeech: A text style control speech corpus with codec language text-to-speech models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:34.257590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:16:31.233871Z digest=sha256:f1ca0bb45ecd0c38ede404583abefaa4d3575615d8632ee73ee897c3b4ec6878

Observation 9d2bd77f-ec90-4f34-8bdb-9bbeb17e9ad7 · outbound

This paper cites Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:34.065052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:16:31.305771Z digest=sha256:340a3d09bf7f4f9b9891ab36f9d57b04b6c701924ff1967ae7478c806e86792f

Observation 4274f1ba-67a4-4902-ae9d-8cb9510965e3 · outbound

This paper cites VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:16:32.936942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:16:31.421450Z digest=sha256:9426fa04323ef6b216293141b562f360eb6e0a48f57ff1707075240049cc4cac

Observation 36e3e659-0c81-48b7-8ddc-5fbf05c4dd55 · outbound

This paper cites Prosody-tts: Improving prosody with masked autoencoder and conditional diffusion model for expressive text-to-speech,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Prosody-tts: Improving prosody with masked autoencoder and conditional diffusion model for expressive text-to-speech,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:33.848661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:16:31.520970Z digest=sha256:774a78334c30feb18fe671d4da314c8f74cd3c5027bcabad0e91ad770c6c5997

Observation 7fc73b80-3143-46a8-8b40-71a6b50a205b · outbound

This paper cites Uniaudio: Towards universal audio generation with large language models,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Uniaudio: Towards universal audio generation with large language models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:33.625603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:16:31.600517Z digest=sha256:5e6872f596e82d300bef7129c138387efcfe2d355415e4bcfb7bf4579633f685

Observation 71595cf2-7d84-4261-bb37-a801c98d8cf4 · outbound

This paper cites Classifier-free diffusion guidance,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Classifier-free diffusion guidance,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.676312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.676312Z digest=sha256:c5eed77fac7a690024cc626af58c6521d272ce9a637e0c7aa43ef5a0ada8d58f

Observation 2c4cdf49-992a-4d22-a6ff-09981ad38bb5 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.762687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.762687Z digest=sha256:e50e962d34ebc43c3aaf9495c96647f2f7261c047d9ef3f0a4477fac8d4f1c53

Observation 827789aa-6c87-4b25-b9ce-1b1bc304e76f · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Librispeech: an asr corpus based on public domain audio books,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.849446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.849446Z digest=sha256:e0687fc57fe4e3ff9c6e0123132fa8378bd8f9b6bc54e592935b730d87257810

Observation fb796082-112f-4c2d-b715-ee251fdd9099 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.921256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.921256Z digest=sha256:8953dce6985b6192526c2118724661452712475ba1c04b5285ac85c31b152587

Observation 6af88a1a-5f9a-498a-b07e-3a8ffabb8bc3 · outbound

This paper cites Dailytalk: Spoken dialogue dataset for conversational text-to-speech,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Dailytalk: Spoken dialogue dataset for conversational text-to-speech,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:33.382842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:16:31.993650Z digest=sha256:aa410d8118e62e19132816460462d92ed616682509142093c3871fd7e01ec9ba

Observation f3069eef-0cf1-4c51-be97-96203ce4197e · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.057240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.057240Z digest=sha256:2227dca12f9b831d17cff05cf8fb3486848efae31cab8fefda663334ec0b68ae

Observation c89fd632-0615-417d-824b-23b60a90d21b · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.166827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.166827Z digest=sha256:e8aa32f42c4f3943b60e556812ef179b2d8e3c93bf110ea19f60944362fb5c70

Observation 1d45114f-5714-4f3e-a082-87767117d8d6 · outbound

This paper cites High Fidelity Neural Audio Compression.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation High Fidelity Neural Audio Compression

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.279195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.279195Z digest=sha256:a1eb64cb9ace58652f67b60b4a8c2668ea276d18b3c1638d819be0312eeda673

Observation 27a2671c-87a4-4c6d-a93e-23daa4105847 · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:33.190288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:16:32.337391Z digest=sha256:046ea0e95c6e19bb4fa92a8420714d505a6ef7d060b6290b47d6bac277789ca6

Observation 975c135b-5556-4155-bad1-0e7301bc06e9 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.422722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.422722Z digest=sha256:e33af9649e9ec0909d0945eb382a633ba97654f12102055f27761aec5cb8e9bb

Observation b4d329d4-8d54-4a05-b4da-311e7867885a · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Wespeaker: A research and production oriented speaker embedding learning toolkit,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.475093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.475093Z digest=sha256:44afa2babb46a3de0dc89f3b58b3880bcc7559509d223cd705ceb99f6a5badd8

Pith citing papers

No inbound Pith citation observations are available.