Pith. sign in

Paper Citation Record · LEDGER

MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 62 inbound Pith citation observations for arXiv:2409.00750.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.00750 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 62 of 62 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:26:03.299001Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:19:49.609343Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 086fc22e-91a8-4ecc-9e4f-0f74191a7385 · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 147

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.545772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:1ce5e4a0f048279e338c3f9928c147cb45be51629917e6611e65733bb1af8bc8

Observation 2e907032-4eca-4045-a25a-8266e682f8a0 · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 217

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:58.141022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:58.141022Z digest=sha256:d9a0b3b0200a202afe387989612a8050fb140c4c4ce3fca555a26185064a3116

Observation 7ac6d73c-4f62-49b0-a77c-5a5d04a92fbf · inbound

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch cites this paper.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:22.378032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:22.378032Z digest=sha256:e846a7141d8ca607f44d472864e6f0a4c8cd9ead4a4d3811e5d2ec0dc0953f7c

Observation ca3a5fb0-7de2-4f11-abd9-268cc17b4ef2 · inbound

EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing cites this paper.

EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T17:27:30.880110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:27:30.880110Z digest=sha256:158050055655a8679c9417a24750e74f8dd2f7bb8800318beb5b8df8b94f8fcf

Observation a67747de-bf49-4491-84c0-2700e7c1cf2e · inbound

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models cites this paper.

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:19:09.552268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T06:19:09.507440Z digest=sha256:e1b0b5fb0b030405a98c8baf709ad07eed1987300fad65a6955a57ed766e54b6

Observation 829c2aa3-ad90-4160-9ef9-f513bb06313b · inbound

Overview of the Amphion Toolkit (v0.2) cites this paper.

Overview of the Amphion Toolkit (v0.2) MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 119

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.990616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.990616Z digest=sha256:23e5fc78a2e2ed8baff8e882903cd9c71a334a17c66dae458b5e4f1ccb1468d1

Observation 637cb4a9-52c4-4d75-abe9-81f326f42521 · inbound

Emotional Face-to-Speech cites this paper.

Emotional Face-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-09T16:50:43.935681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:50:43.935681Z digest=sha256:7374f618a4408ad75c51c47c96085d80b2aea56cea5627c8927755f18f73021e

Observation f4d8a91e-e789-4371-b7d5-103e43f7811c · inbound

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training cites this paper.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.783977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.783977Z digest=sha256:98c740953785bb1f1fd11a597dc02ae6ecf06ba80404c116652c9414c589b5df

Observation 80fd235d-9dec-42a2-8da9-4a87396ebd46 · inbound

FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing cites this paper.

FlowDubber: Movie Dubbing with LLM-based Semantic-aware Learning and Flow Matching based Voice Enhancing MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:03.299001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:03.299001Z digest=sha256:14e46b19c51ece4bf00dcc89794c1444d4961d1e41ce01289c53adb8ac3855a1

Observation da8156d4-ecc8-44b2-a160-ff66a880bfa4 · inbound

SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation cites this paper.

SepALM: Audio Language Models Are Error Correctors for Robust Speech Separation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:00:56.877406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:00:56.877406Z digest=sha256:4a8602f196d3c7ee8086aa1c76306602f5574f8ec650420aa42f5a6c9cb31a63

Observation c26f3345-058c-47a6-927e-1fc6efdd23e9 · inbound

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations cites this paper.

Teochew-Wild: The First In-the-wild Teochew Dataset with Orthographic Annotations MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:54.758777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:54.758777Z digest=sha256:7007a0d7406b6e023503d8315e513a52c38a0df52893fc186549245151025fc8

Observation 4648c6c9-9105-46a3-9050-85ee6376b56d · inbound

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech cites this paper.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.242589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.242589Z digest=sha256:1596c27fadc45e2387abe22710a3e83df2d832f5988e36e6413c020b90a262f6

Observation 5e71f19e-d50f-4ae4-8075-535d96840a6a · inbound

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder cites this paper.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.327211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.327211Z digest=sha256:75349af4f06eb9e2fd7259d4a01ee52309ed2a62e2ee436594252bb4d1542bfd

Observation 323ba844-defb-424d-9b4d-76a2e892d01f · inbound

A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model cites this paper.

A Semantic Information-based Hierarchical Speech Enhancement Method Using Factorized Codec and Diffusion Model MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:14:08.477692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:14:08.477692Z digest=sha256:82e696bf709715673f5165dcf8d3cd60112bab376e47eb5011a21ef37889ec4a

Observation 1e70ec91-873c-42c2-894f-28e9d834de26 · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:25.504594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:f33b777098e88ae3ee08af59c90a0c4938f2e33680218c9aefc6227e9650db8d

Observation 03333264-800a-4980-89ee-16fd2c942e25 · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:51.135367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:51.135367Z digest=sha256:97cf97acf481be00501d0bfbf6594e0a85c4b22d34fa310dc5326d9f0fd9398f

Observation 9ba44ab7-e2ce-4b97-b982-f296b1bfdde9 · inbound

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling cites this paper.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:19.616970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:19.616970Z digest=sha256:9fa960639475a08f16fee1cf0602d6cda88fbedf1de7fe9569a3553cfa2c4ec7

Observation 88db68ac-66a9-450a-8a21-beb069a6fa7f · inbound

Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages cites this paper.

Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:54:23.558229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:54:23.558229Z digest=sha256:103e6f1b2efca4f2c5326239e93ee1613b30706f3476dda8db78b460409a22be

Observation ee274426-26c7-41a7-83ea-6e304359fe61 · inbound

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation cites this paper.

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:29.941725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:29.941725Z digest=sha256:35bb19e404ceacf8013b687b24b554f0815cc26c33c3a41e245aa131273f768d

Observation 22e0c373-417c-4506-8d24-986b3f37aa1c · inbound

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification cites this paper.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:39.416451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:39.416451Z digest=sha256:ac5c77cc352f9bd8ae500ccf7b6dd1ec36ed8f18b95e06cca2f8972432f5011d

Observation 663c1c04-c3e4-47e7-9303-e95d307ed5f0 · inbound

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis cites this paper.

A Multi-Stage Framework for Multimodal Controllable Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:43:50.332197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:43:50.332197Z digest=sha256:de676412c3f2f46571a880ab3beb9f18bac4671ced79a4807a280c984041d4a0

Observation 04e4ef9c-3e63-4826-b1d8-4ba7a760e217 · inbound

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech cites this paper.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.231724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.231724Z digest=sha256:f547fb87361e1a1a23e8ffe2bb2d777443d498e5059f8f9c5d95ff63fadc8c28

Observation 824d47ee-60fe-4859-b56e-c26291a5af5e · inbound

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges cites this paper.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:51.703103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:51.703103Z digest=sha256:bf4aad38f054df4412f8951bc8ee0e8c8c623e5d5f8ba7eb66939c0c9f6d2058

Observation caa92c4b-ae16-48b0-9694-c7636e770379 · inbound

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching cites this paper.

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:32:03.632822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T04:29:41.285194Z digest=sha256:737de93c49e0d7045322279fc5e368708d524b7f53f90af29d1afa8890a7b81b

Observation a32f064e-3d61-4988-8f59-e4bee014f2fa · inbound

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations cites this paper.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.368580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.368580Z digest=sha256:e37bde89f1e39024245d947f814ba2603d46dda02a265b458f0c2dae36b75b8c

Observation b3cbe4af-0b4a-4535-b907-5547b8ff3d60 · inbound

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis cites this paper.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.576312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.576312Z digest=sha256:ee266f3869fabaf2ab470cd348b882b80d3b0472f686e8c2c8721744496f6b35

Observation 63183282-1e75-44c9-94dd-9a9f6e44dfa2 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:51.066886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:a4ca0ef7f7fcc4cf89b5ffc98cde8858f99a5b555db7859314a85415b70861b8

Observation 752c8ba0-a959-4247-8d1e-c3737aee4cf3 · inbound

Adaptive Duration Model for Text Speech Alignment cites this paper.

Adaptive Duration Model for Text Speech Alignment MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:09.403232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:09.403232Z digest=sha256:1341097f1751933657293a2ed649d54c1dc53fa66c95e9cf029791ae4b9a3420

Observation cd815554-0fb7-4877-a0b9-f2406c475e7b · inbound

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation cites this paper.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.080956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.080956Z digest=sha256:0df1cf46e9129c144d9b97b5ad76c57996f6b8771d30f52caa1471b06deb6921

Observation 46ccd477-b2ae-4e42-ac81-b7a3e019362b · inbound

Entropy-based Coarse and Compressed Semantic Speech Representation Learning cites this paper.

Entropy-based Coarse and Compressed Semantic Speech Representation Learning MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T13:36:06.639779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:36:06.639779Z digest=sha256:bbceb15ba5578d157401be1ac3a1ab6ecc543852edc9f0d801797ff5738c3ae6

Observation 6d8e8d91-3e33-49eb-8da4-a442370d71d5 · inbound

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot cites this paper.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.427743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.427743Z digest=sha256:fe7ec5b3a55cb6c854b970e0f5a2053f2eed9b79d27b72a62b5eb7d9ceef5984

Observation 751bb11e-20cb-4fe1-a0c7-cc73a5d917cc · inbound

Audio Deepfake Verification cites this paper.

Audio Deepfake Verification MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T16:12:07.923921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:12:07.923921Z digest=sha256:5c4323adbf4ea0683437bb2c29ca68b626046e763a93ee00032fedabb05c8773

Observation 93e93911-012d-41c8-95b2-a035cc226d4f · inbound

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration cites this paper.

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T19:11:59.598937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:11:59.598937Z digest=sha256:d02a392a80f941e6ec6c0c6f67f86fffec0d8e22aae4f908ea121003574a3486

Observation f634f89e-bfc2-4b66-8aa0-d610f9a94a59 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:36.628124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:36.628124Z digest=sha256:8cb131c6d677e5faf833db91044e1d635ee32f41ed3988541919926e8892142f

Observation 45072d0a-4ed1-4e45-8486-d449f290cad2 · inbound

TokenChain: A Discrete Speech Chain via Semantic Token Modeling cites this paper.

TokenChain: A Discrete Speech Chain via Semantic Token Modeling MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:16:09.506545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T09:14:58.542628Z digest=sha256:25e0955cc06913c4ec99071c36beea15c628461aba1e9779186f0d005172604e

Observation 8f2ccafa-0bd7-47fd-bd41-f2e79f8858f2 · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:24:56.159190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:f82cc0a2e8b9f585fe1ee834a6b080ce76b3320030ea7d75b9478c2eb832159d

Observation be168de6-39a7-449c-9b49-70883c5b8409 · inbound

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck cites this paper.

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:50:51.681996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:51:38.030059Z digest=sha256:159b95e8905bab23aed0afe157fc602cabf1809e4ff8cc76cdb12edb0eca326c

Observation b7b74c95-bbe2-4c41-8ceb-22ebd5ff1b46 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 110

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.196874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:bfb5757ebc4354a6ebfbbed9d93b96237a7409b9d007939fdb41c479996b11ec

Observation fae34acb-80fa-4cd5-a6f1-5152c13aca10 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 101

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:794bb008fa7b0ce36a99a9de3461b6cca6f78c1c93f551d16ea5332954cdd72b

Observation e4a5a0c1-14fd-4266-8ec7-75aa23856e8f · inbound

MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora cites this paper.

MimicLM: Zero-Shot Voice Imitation through Autoregressive Modeling of Pseudo-Parallel Speech Corpora MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:26:01.400475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T14:57:07.894455Z digest=sha256:0eb5dfcc40ffd51e14c37a0426dc6ad7ebf503d376eb2d87aead120cedacfaf2

Observation cd698a02-fa82-4e54-85bc-f48c8cc23020 · inbound

Hierarchical Codec Diffusion for Video-to-Speech Generation cites this paper.

Hierarchical Codec Diffusion for Video-to-Speech Generation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:12:26.301870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T08:12:01.260833Z digest=sha256:41b6ffcffe3c6264bffe3d71ec6afff4038f9b5fcd3050eb591f4ed7052befa7

Observation 9548e9dd-0c86-4d8a-8d3e-e1dd3b2ecd07 · inbound

Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation cites this paper.

Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T01:04:50.114051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T01:01:06.094276Z digest=sha256:c7c919366f5662cb3ece446f840367492082a5d19c0df691a9ed9d14edb7f698

Observation 2587db95-d30e-4031-941a-ee853a6e8982 · inbound

UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions cites this paper.

UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:21:12.858379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T09:18:04.870414Z digest=sha256:2916eb5d59a0295a765ab2fe01a3b42a84a9bec488bb03bb78c126aeb3443017

Observation a8e0ff9b-a39d-4744-837e-4ce17ed55243 · inbound

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling cites this paper.

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:07:00.271270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-13T01:04:54.506749Z digest=sha256:95a7604145aece1ff7282e60068a2f233b09ec9b8ae72ae140bb718b2fca7c95

Observation b8583be2-aef0-4be5-98d7-8b637f38461b · inbound

Taming Audio VAEs via Target-KL Regularization cites this paper.

Taming Audio VAEs via Target-KL Regularization MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:53:23.231264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:45331bb24b9e5088938119adb5ee65ac7344d52817b9ed6f54b32bbf4abff003

Observation d84b2da4-4d8a-4534-98c4-d35e80872289 · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:13.128977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:5c94b1ec0a747ecf10477ca144b01f86f4f7044c67209db345bd07b6892127e2

Observation 878f64a7-6f99-48d7-b9e5-f790dcad259a · inbound

SegTune: Structured and Fine-Grained Control for Song Generation cites this paper.

SegTune: Structured and Fine-Grained Control for Song Generation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:36:14.912924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T16:46:40.934795Z digest=sha256:33c6b204b48a53d302f5467ee65adc36976db449e3d43ff065c3b113ed5d5646

Observation 2ee16ea4-22a0-4f0d-8c81-d09cd0a93592 · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.738763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:96bffaeded7ca5db6b660f3af8eb140d430ff5ee0009b4c6649029119e9dfa85

Observation 52ae7175-3f4a-40bf-9a0c-b0535720a546 · inbound

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech cites this paper.

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:07:35.869017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T15:27:01.142747Z digest=sha256:f4222cc70c86e428433e01653c88fa2907b7b58b270492aeaae3a626f76ebef3

Observation 372151ba-879e-49d0-939c-093574a86af6 · inbound

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation cites this paper.

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:35.149517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T15:15:10.060770Z digest=sha256:2b067a350475fa2ff252c317befca1f1bc10aa23023f46850368919988323385

Observation 3adf4026-52b5-4038-87fa-1fee4aa4c4b5 · inbound

End-to-End Training for Discrete Token LLM based TTS System cites this paper.

End-to-End Training for Discrete Token LLM based TTS System MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.897423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T15:22:06.893507Z digest=sha256:277b6609b0eaacc4376585361914f0fe5dd7ff84c3b7fcdee73a52732b9c76ea

Observation 026bcbb9-2253-4d3a-bbbf-3a0715c81c4c · inbound

An Evaluation Framework for Text-to-Speech Voice Reconstruction cites this paper.

An Evaluation Framework for Text-to-Speech Voice Reconstruction MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:39:38.710628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T13:16:05.358573Z digest=sha256:b9d80fff65dce1edf110c7ac4c8d95280efe84def97e59e61444e7318a8d1331

Observation d10e45ad-c73b-40fe-90c6-c06deea7f08e · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:19:49.610738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T07:02:36.499424Z digest=sha256:d76000821ee6700956b61ac0739ba7ebb9062038665d609e84237ddac9cd556d

Observation c04579a6-0ded-4ff9-907c-34885ad06637 · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T12:44:20.831164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:44:20.831164Z digest=sha256:184ad8966975b7c15ede7679a38b913ae8b72948fe8d5179a1517a01e26329b8

Observation 044e4142-2a3a-43d1-8098-8415a551a242 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 202

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.141550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:1a5a9efbe7cf85119f7cd5d55f05065a97ec97a6ff2cf69970c719a9a87e1553

Observation 418d0736-613b-4ee4-98d4-f3e77d558404 · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.856520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:c41430cd2c77e5fab2e045a1ee68833abe636c6cb9076ce606eae334d87abaac

Observation 60ab2a50-c696-4b5d-905b-9e54efd22029 · inbound

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis cites this paper.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:11.847608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:11.847608Z digest=sha256:56f57b2d7a1dad8ddef375232af8312631bd409cd2a3997af13e0e5d202a9b2d

Observation 0ed90ffa-6d2f-40e0-b4aa-a8a78347f78c · inbound

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis cites this paper.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:37.876645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:37.876645Z digest=sha256:4d5dab32dcfd98c7623faea60e02aaf5f85b49d072cf7f404765d4f096ef7e58

Observation 4423d7a2-cb63-4f14-a32e-0d9655df8656 · inbound

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation cites this paper.

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 158

Resolution
unresolved
no resolver link, observed 2026-08-10T04:20:42.329390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:20:42.329390Z digest=sha256:83ed8c7f2966edc970d7df2e331face4d22279e85cd12c5ff264e1f42aa547a3

Observation 4dc0ab8f-49f9-4847-84ec-fe1532ef48a4 · inbound

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis cites this paper.

CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:10:40.826976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:10:40.826976Z digest=sha256:2a312fa476238feef24858ee326e15b448817ba575aaa8739db376345eaad4d8

Observation d2eecb1a-ea0b-4bf8-93c3-ff37f5c41f39 · inbound

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder cites this paper.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.752186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.752186Z digest=sha256:ba76d98e4684ffcab9394b14cf8c239dc3398be399f74d9491f0992054f4682c

Observation ae40bb28-6d73-43a9-a72c-fa9cccee1085 · inbound

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization cites this paper.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.990844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.990844Z digest=sha256:aa6b5915af7015dce49809228d5f459d365f71b7f8f22823b4871931010ed3cd