Pith. sign in

Paper Citation Record · LEDGER

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech

As of 20 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2505.05159.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05159 v3

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:17:25.271659Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a9490f9d-24f5-4afa-8103-79b7378c26f1 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.024367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.024367Z digest=sha256:3586ba22ea5c484255597955bb96c6b6a3ba07c817e6330535c1957e6271147d

Observation 7f315e98-c5e8-4d69-8a82-eacaca25f95c · outbound

This paper cites Tyers, and Gregor Weber.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Tyers, and Gregor Weber

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:26.195369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.028956Z digest=sha256:00e7faec82abff1dc33f63ce03a8c730e53cff350d11e8bcb255de960e62fa0c

Observation 47971306-ecbd-4448-9537-17ea147d8861 · outbound

This paper cites Better speech synthesis through scaling.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Better speech synthesis through scaling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.033320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.033320Z digest=sha256:ffa326aa5d46fb7699a8f86736ae0e86de719b7c3a50aa803a8b3d664e9551b3

Observation 55329d24-ee78-40e2-8ad3-3820a221f0b7 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.037475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.037475Z digest=sha256:548944050b51119d9c0d59eeebcf347d9cecbe79646eb29cb65af8f168a68a19

Observation 78e2791f-d7a9-4f52-8f5d-0c02a4e264a9 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:26.177988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.041341Z digest=sha256:d57eda337ae2a51796c76117e63e20c7a6328bc72d000fca5f0ef116fee14d10

Observation b28a1746-2ab3-4a8d-8d68-835476ebdca4 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:26.166509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.045495Z digest=sha256:db40ab1884adc034a17979968fb1934761ab0bf984b4500f0dc708f2f3bef690

Observation 07bbf205-96f3-4b14-a5ba-b6bd3c6cded2 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.053367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.053367Z digest=sha256:25f8955ec6280c4b07c9ed906d6c221986ecb520df0270a744ac9eb59c40a053

Observation cc148e6e-d9ba-4144-b0b9-74d09a006111 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:26.154640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.057095Z digest=sha256:74100ed921ca471c623542f2e32abfb34fd787c1f6c2495b4a00788b551baa4b

Observation 501070d8-e3ec-4521-bdd5-5ab4fd1102f4 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.060752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.060752Z digest=sha256:2a7960ff6eab4f27dd2c6f1ea63c2fd84c4bf4b0ab25c131981ec1101236047a

Observation 5289587f-1e44-4194-833f-80fb612610e0 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.065429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.065429Z digest=sha256:64db592a96f878db82b56a0a3b04fe6162639972681c0a140f144cc430f13020

Observation 3a9bcbdf-5c84-4293-9fd9-fc4655b1a735 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:26.136352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.069143Z digest=sha256:04e15474e6a2e8f9f2baf0f91ea70d938976f770b96fcf605941cfd73e3e2734

Observation 94ffb4b5-e451-41a1-8aa4-0fe2cf4d3bc9 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.072873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.072873Z digest=sha256:7a83216b2d8f28fde885bdd83b8ffcf64c40c6a952e4209b32a089288bde878f

Observation a9c77cef-4f8d-4e85-8431-c71a56cf7ca0 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.076827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.076827Z digest=sha256:1d09c02b03bb2db69173bf078f6f02d31d58e996ba8e8278342b25563b387f3e

Observation 233f2909-d7df-4113-8830-85c23646da5b · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:26.124791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.081403Z digest=sha256:d5d76e9a0f558fe7da95010ed8172cdd6a437f66a6a405235ca30510b0d57060

Observation 79d3f9fd-10b0-49fd-92b0-8d599d4fddd7 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:26.112869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.085052Z digest=sha256:647a6f44e094e77e532b41e6aef4f4627eb99b12dc30471f1a83cafca1ea3770

Observation 9cf70719-1813-41ab-a409-638b3458340c · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:26.099757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.088569Z digest=sha256:1590a13e6f0024a45d2aa7961a6412fb3f03fc127f5fb5c530f9d9bb95940018

Observation 16276d46-c644-45f0-995b-86a7fe915482 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:26.086373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.092773Z digest=sha256:d7f5881d1a0fba0a3a5cb6ad00bf62771fca248e92ef8070184bce2c0463a826

Observation 42c41b85-8d58-40f3-8f85-d0f2ee1bd6f1 · outbound

This paper cites Classifier-Free Diffusion Guidance.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Classifier-Free Diffusion Guidance

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.096810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.096810Z digest=sha256:c787bfc742cb9f6607cad3059848614c14897d82df925073b386855a3a66d5f0

Observation 126324db-0e16-4baa-9d02-9df7014944a0 · outbound

This paper cites Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.100800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.100800Z digest=sha256:451b6de7c1ef2ae902d0a30ef5a7ec91ce0bee4c89afbc87633d6cb7bc6cf0cb

Observation 9ab3db31-3229-4947-a6d6-5a7685a65f82 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:26.074415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.104784Z digest=sha256:1796b2ba88eb521855fbfd04b99924a9446501937ffc6f049de909a8c10c0aa9

Observation c7451e47-a31a-433c-ad4b-0098950e221c · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:26.062639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.108479Z digest=sha256:cf3cd63a31e4fa1fec0748059929caf29ba6d447ac3bd60ce99485c33806b382

Observation 2326d4fb-f06a-423d-8975-28e7f84d62ad · outbound

This paper cites Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.116474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.116474Z digest=sha256:88bea7edfdebbf3ec09659dd7c4aad3f06a48125af76dd796e6f12a072712d16

Observation d172dc6c-58ef-4fc3-ad3a-b58e8d402168 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:26.050879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.120655Z digest=sha256:73b966896bea18d9def0943d777f2b6e450c059658f77c0f19f7d9b50b2131da

Observation faff6970-8737-4e8f-971f-d98d7235c3d8 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.129432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.129432Z digest=sha256:b3675d19ae2d3e1c57e9d47ecd5b8b86b215b9b2d0c9cddb07652e858ccf7fa0

Observation ea2cfb0e-81bb-4b7f-8b79-2a9fccae81e6 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:26.009314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.137141Z digest=sha256:0d650e2f825d054d54f1f8d27c73bd4272c91ffbbfd74e487c4c7cdd8b46ab29

Observation 5263a8f7-cd24-4209-afbe-46c86ec8e671 · outbound

This paper cites In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:26.039710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.125194Z digest=sha256:c40d2daf729b8897f6fcc5cea9be41c10a424fcf620cfbac642c760cb5f5af13

Observation 43d65f4b-9a9f-4c18-a096-d017b6dcdb45 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:25.984811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.144914Z digest=sha256:2b33aca05f48997ee99a46bc383171b12dc24900a4d525bd69ec64b3f9974e2f

Observation 8e81e9a7-447d-4a34-bc5f-5c61ff2f3b17 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.153279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.153279Z digest=sha256:644578920e3200ce38cfea401ba6f9f10840ffb5136690a3b6769784d6605e35

Observation c9f39c23-90da-4117-8f88-343804f73a6f · outbound

This paper cites Continuous Speech Tokenizer in Text To Speech.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Continuous Speech Tokenizer in Text To Speech

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.161455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.161455Z digest=sha256:78f1da634401a3187dbbc76b59ed8d630bbe1ec9ced5de7b3609afb528b63af0

Observation b1f8f61f-8a88-4465-bca9-755d8118a00e · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:25.997556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.140962Z digest=sha256:0dbdf78f735249baca8d03c1c21067ec05e4074dc6b99f799e7922174f38a14e

Observation f7d54ad1-ba27-4e84-86e2-55e0312b3758 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:25.954244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.171419Z digest=sha256:4a078185ddf25647c9ddb826ba91d2d3f1f20c717509639b319f3763992335d1

Observation c6759a81-8496-4955-9aaa-4033dba87444 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.148966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.148966Z digest=sha256:4c70b67c9702c9dcd310a727ccbbdb2c41083c65761bde298f5ccbbffd106e7f

Observation 368429bc-1add-4ea9-82b9-6e8cde7f4364 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:25.943842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.180171Z digest=sha256:20b4800103aa8176748ea49ae899768865aa17e31deff6ea368594423e0194ab

Observation 78a5465c-2da6-4ddd-8b59-3666eb41d44b · outbound

This paper cites In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:25.965733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.157451Z digest=sha256:f9dd5e639826d26e7e6c88a3a8c36ab246c99bcd9fbdd8d48782474d3886ae79

Observation 177df6cd-5081-4509-bda9-265e20a1f142 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:25.922752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.192471Z digest=sha256:64285a6b5146ebcee6267aee91fdea8856d211db1974a87dc0e00f74994c194e

Observation e8ed7c07-7ba5-4896-856a-ab21e0eb11b6 · outbound

This paper cites Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.165930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.165930Z digest=sha256:60029425cfd2ecc52f9ca5947d931819b65b4b31ee384447911e0be586d3d6de

Observation c3808849-4f14-4525-b6a1-07a3e14832f9 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:25.897403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.202038Z digest=sha256:fd6ab9f75b4e1a993943eb88a28038bf08e8da087105dad458254c75fb8986ac

Observation a5a9018b-7524-49c7-9b88-4d334c05eeb7 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.176165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.176165Z digest=sha256:84a2c11a48e20a9f7e8c7219a0eb19a1453688e8d15b5e8ca4392145d6bd3cef

Observation a72bcb5e-2b1c-4e37-aedb-830bce338cf9 · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Manning, Stefano Ermon, and Chelsea Finn

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:17:25.870707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.210857Z digest=sha256:8261b235d4c1b5ff71a34fc2546daa6920c110406c92926b90aa945358e948bc

Observation e041d6ad-37a2-4169-961b-87ea194fce8b · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:25.932905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.184235Z digest=sha256:ae7a44573b58b676bca047173f8d1f8b1e5ab837875601740b56c5749f253013

Observation 9d458479-2df3-4c29-b479-039af330aee8 · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Autoregressive Speech Synthesis without Vector Quantization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.188076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.188076Z digest=sha256:b985558a1467426d4368f78338719d57a4f93e1bdb882436a778a21691d07536

Observation 9e53aeb0-b76c-4c91-bd7f-d23066248204 · outbound

This paper cites Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, R.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, R

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.222371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.222371Z digest=sha256:5f4f31f24ab6e59cf122a1a99c086f323f024dd89796b73b6e0705025ccee873

Observation 3e900965-a3e7-4b76-a2d3-7606abb7a2ef · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:25.911672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.197273Z digest=sha256:9815f2be7aa62f511e9cbb47f664902ce1a52e10edfdcaeaf07af5213e78109a

Observation 8c2eba9d-6f05-4273-924b-a1a22cd3e20e · outbound

This paper cites Preference Alignment Improves Language Model-Based TTS.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Preference Alignment Improves Language Model-Based TTS

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.230214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.230214Z digest=sha256:a5820fd62399a767edd7459badfc2e6d127f09270292903fc55f8c629cbc8b44

Observation c485c751-a2d2-4e37-98f4-2daee884d681 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:25.884877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.206939Z digest=sha256:e88317863245b4e53509f14d42b54479ebeac48bd87f13f81da18452ac4e2b5f

Observation 04cd8926-2c33-4fcb-a368-f4a59a78ed46 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:25.827936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.238834Z digest=sha256:cf3c63eda795457a95c0c29e2ace1fb4e03ea4cabebef2fb86c0b9de4fe95c74

Observation da2a1e4f-cc12-4a8b-8a27-c1e10cac4c53 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:25.857496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.214679Z digest=sha256:f6dc6a8082cc21d21a1d2e1cd5c1d73a5e49b6fa38c8ad824c6ee8a0c7a8cb6d

Observation 872fd3d6-1e52-4de4-80c1-1ba13aab6c9b · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:25.846310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.218619Z digest=sha256:0fd2c2172956d4b81a2637a61e35b24cf200d3273f2eed7c1d5132ac47ed91ef

Observation e721eac9-0492-4384-9926-c7f0c9d8419f · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.251071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.251071Z digest=sha256:cd8fd3043ee6ecae262df359e67fad95f986b0d4dcafe94a4d9201381b328f95

Observation 39638f6d-39c3-40b7-8ef1-1c49a73027a5 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.226168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.226168Z digest=sha256:1b57dca99ce2492c98da4b4aa8fd06a75f34919e3754a1f4de3a6710d45f0de4

Observation cf870984-9d35-4a93-b939-a39f0cf600ba · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:25.803893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.259252Z digest=sha256:6cf8606bd3005fc81f098e130b262a11eb93b26e47028b8380158127cd4b42b2

Observation 83dbe4eb-e066-4bc6-8f35-6d34553c75a0 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.234328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.234328Z digest=sha256:6c353f9d96bcc44db1141f02fb5c6aaa0ec98b0b4c01000cd0631d45fc6711e9

Observation 4b397100-866b-42f1-8035-b589747f1c33 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:25.776375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.267769Z digest=sha256:b22c71f9c32034cb8fbf827eb0f83932b0fd11cbfc2452f56562cc53afaf5d6b

Observation 4648c6c9-9105-46a3-9050-85ee6376b56d · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.242589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.242589Z digest=sha256:1596c27fadc45e2387abe22710a3e83df2d832f5988e36e6413c020b90a262f6

Observation e7e3735f-62ec-4eb1-b101-c61ca7261c0c · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.246725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.246725Z digest=sha256:654675dcad4a1bbd3a76f7f5ab0a96def2b559f81283addb72e8266da0442e49

Observation d4886e0b-8877-4bef-a76a-e7d8311bd623 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:25.816676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.255469Z digest=sha256:b686eb343ba16766e2a62570ddd830074202577208e9fc71c6b4e86689e0a182

Observation c37abb75-f70d-4d98-8de2-3918b7c75c1e · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:25.790163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.263387Z digest=sha256:36fa129d3fe7afbfa924f396b308d84234af956650be61d645908011f5fc0da3

Observation b07c6bc7-f60f-4a27-b419-5d77b7a97369 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.271659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.271659Z digest=sha256:91502dbfcc846febbc4b9ac1e36b8914998027f238321de8021769634f715c09

Observation 64be8acc-0ce3-4cdd-b726-5f6024bece62 · outbound

This paper cites an unresolved cited work.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:17:26.020959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T23:17:25.133484Z digest=sha256:a764b788b490b59602b8f01381ca2f5c7409047c12f286793244ffd631dd0a00

Observation 629f9cd0-9994-4a76-8b3c-c4bdabb64aa6 · outbound

This paper cites Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.049265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.049265Z digest=sha256:7fa7202935d4e757b84049b5c5230dded2c705c9ea5711cd6930b377612a67a8

Observation 2e77d247-b361-440c-8a45-8d12cfda49e6 · outbound

This paper cites MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis.

FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:25.112300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:25.112300Z digest=sha256:2f21b4c2e5c6bac599718aeee2bb37dbf851f1f990b66fed1b109147a8085716

Pith citing papers

No inbound Pith citation observations are available.