Pith. sign in

Paper Citation Record · LEDGER

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation

As of 22 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2506.02997.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02997 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:16:32.475093Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 98ee0846-7ee5-4696-a363-c8518914cb4d · outbound

This paper cites Prompttts: Controllable text-to-speech with text descriptions,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Prompttts: Controllable text-to-speech with text descriptions,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:34.432358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:16:31.033516Z digest=sha256:f50ef79ce445e2ff013a0c049963fb228f253ec4c652f9d8b7939b9e0b87324a

Observation 43e8d4ab-718e-47cb-bf3e-fabe5fd92d5d · outbound

This paper cites PromptTTS 2: Describing and Generating Voices with Text Prompt.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation PromptTTS 2: Describing and Generating Voices with Text Prompt

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.114971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.114971Z digest=sha256:c5df5ea5d9cd65cd5d1498d5f32b959c6f9cd7a697c2a075356657a8d7bf7c08

Observation 01f03004-b063-4a86-81c5-1c4eb5d14d2c · outbound

This paper cites Textrolspeech: A text style control speech corpus with codec language text-to-speech models,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Textrolspeech: A text style control speech corpus with codec language text-to-speech models,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:34.257590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:16:31.233871Z digest=sha256:6cd46e8426521585d3f458443a79e6d91300e84784436b86ccdbe42c2bbffcac

Observation 9d2bd77f-ec90-4f34-8bdb-9bbeb17e9ad7 · outbound

This paper cites Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Instructtts: Modelling expressive tts in discrete latent space with natural language style prompt,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:34.065052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:16:31.305771Z digest=sha256:2b723e3feee6e23c407244a6f2c9bfbc3d91688b4888c039c9e69ea9e3f2d6a9

Observation 4274f1ba-67a4-4902-ae9d-8cb9510965e3 · outbound

This paper cites VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation VoxInstruct: Expressive Human Instruction-to-Speech Generation with Unified Multilingual Codec Language Modelling

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:16:32.936942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:16:31.421450Z digest=sha256:dbe07ff0ad3ff8d884f6362e821f9162735900c2d6028ea607d1b73216765db5

Observation 36e3e659-0c81-48b7-8ddc-5fbf05c4dd55 · outbound

This paper cites Prosody-tts: Improving prosody with masked autoencoder and conditional diffusion model for expressive text-to-speech,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Prosody-tts: Improving prosody with masked autoencoder and conditional diffusion model for expressive text-to-speech,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:33.848661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:16:31.520970Z digest=sha256:8a6a6fce047594dba12a7ddd48a4da037d9923d6b1deacb141b609ff7bfb70fb

Observation 7fc73b80-3143-46a8-8b40-71a6b50a205b · outbound

This paper cites Uniaudio: Towards universal audio generation with large language models,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Uniaudio: Towards universal audio generation with large language models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:33.625603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:16:31.600517Z digest=sha256:3714e03a2a5b38f66f27102d59fc92a88269ae142de482d18472ef148685dd7c

Observation 71595cf2-7d84-4261-bb37-a801c98d8cf4 · outbound

This paper cites Classifier-free diffusion guidance,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Classifier-free diffusion guidance,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.676312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.676312Z digest=sha256:ddcb1d02116b4828c0314b08795151fedf3b3aec7df2afbfa7f56caf8af24d00

Observation 2c4cdf49-992a-4d22-a6ff-09981ad38bb5 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.762687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.762687Z digest=sha256:1886c6edae386a72f876cd91c92c3c326149d399423cb863b6cd104b858adce2

Observation 827789aa-6c87-4b25-b9ce-1b1bc304e76f · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Librispeech: an asr corpus based on public domain audio books,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.849446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.849446Z digest=sha256:08397097c81a3ef7a0bffa06b2ecb65196c6e18ab387084f07c7093657f186c2

Observation fb796082-112f-4c2d-b715-ee251fdd9099 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:31.921256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:31.921256Z digest=sha256:559aa962b4f7dc8fa9273e9d44b5739e13bcce16bfd94bb83c3944b39140b0b8

Observation 6af88a1a-5f9a-498a-b07e-3a8ffabb8bc3 · outbound

This paper cites Dailytalk: Spoken dialogue dataset for conversational text-to-speech,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Dailytalk: Spoken dialogue dataset for conversational text-to-speech,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:33.382842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:16:31.993650Z digest=sha256:445a981c3fc5050288552b6b82da487e86736212e4f232982e6ffed6fe4373ba

Observation f3069eef-0cf1-4c51-be97-96203ce4197e · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.057240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.057240Z digest=sha256:2865869bf66ef366f47e0457f5a827616ba6cb9b14ac410f62a76a17b0d6b6ae

Observation c89fd632-0615-417d-824b-23b60a90d21b · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.166827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.166827Z digest=sha256:0c96bedc6b5686ee5770cd3865a092e3b5e7036a2644032f2bc8946a23b1ee1e

Observation 1d45114f-5714-4f3e-a082-87767117d8d6 · outbound

This paper cites High Fidelity Neural Audio Compression.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation High Fidelity Neural Audio Compression

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.279195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.279195Z digest=sha256:73573d9bd4d6c7aa9ad38ae10cacb05264df516409f9b624aac6cbaf8b0aa24a

Observation 27a2671c-87a4-4c6d-a93e-23daa4105847 · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:33.190288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T11:16:32.337391Z digest=sha256:e1c32e3e987821aace11b7e88621cd9170398f938596da11c1c8fd29b81746a1

Observation 975c135b-5556-4155-bad1-0e7301bc06e9 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.422722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.422722Z digest=sha256:0e9fca9098c2f47ded3f00b02488f79beb95895771b9bfd94bbb9b448fd90e7f

Observation b4d329d4-8d54-4a05-b4da-311e7867885a · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit,.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Wespeaker: A research and production oriented speaker embedding learning toolkit,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.475093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.475093Z digest=sha256:012af1a969e8d32bbf7d415a4dc18447b591f1bae47b5190f0f18265455638cb

Pith citing papers

No inbound Pith citation observations are available.