Pith. sign in

Paper Citation Record · LEDGER

Better speech synthesis through scaling

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2305.07243.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.07243 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:30:26.989835Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:10:07.252358Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fbece7d4-c200-461f-a8ae-be65655fed70 · inbound

MLAAD: The Multi-Language Audio Anti-Spoofing Dataset cites this paper.

MLAAD: The Multi-Language Audio Anti-Spoofing Dataset Better speech synthesis through scaling

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:36:00.894144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-24T04:34:59.205522Z digest=sha256:368dbb5a03d7a2f919b9621a5de6f9dfbe51723af5593b2a724bec3d5f67466c

Observation 36c79e26-f01d-4b1e-954a-a6c8070fc732 · inbound

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models cites this paper.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Better speech synthesis through scaling

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.369560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:e153afeb0453d974d5cb66f66d9a069aedd00299e72eac5f95e5a4d5b7daed8e

Observation 4e19c775-5283-49ee-b401-6287a4341f94 · inbound

GenVC: Self-Supervised Zero-Shot Voice Conversion cites this paper.

GenVC: Self-Supervised Zero-Shot Voice Conversion Better speech synthesis through scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T22:30:26.989835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:30:26.989835Z digest=sha256:0b17f799ce17599c3ee0d02ec776bc74b2c0b457685b656fe2dc282d4e2919f3

Observation ecf8f860-9809-4c0e-8dfd-c0805f94c1a2 · inbound

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System cites this paper.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Better speech synthesis through scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.542381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.542381Z digest=sha256:d38ebb6be2214164eafa8e79778591a203d99e8abdd1c99b0fea901585c85a61

Observation 54dc202b-f995-44b5-9425-fc2300fe4d46 · inbound

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement cites this paper.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement Better speech synthesis through scaling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.165401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.165401Z digest=sha256:1423561de31df12e1209750d0a200e7bb98ddcd46d9b4940d3bf9ee0850a1437

Observation 6c64258b-9b03-4a75-b2c2-c5ce3c09c76a · inbound

DeePen: Penetration Testing for Audio Deepfake Detection cites this paper.

DeePen: Penetration Testing for Audio Deepfake Detection Better speech synthesis through scaling

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:45:19.482982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T02:43:04.745094Z digest=sha256:238d62cd65d45b08cdab53750ea3343e0bbad39fcb94202de1bcdc847fcfd008

Observation 4e51cc1d-d6a3-47a1-9995-9a7b2bd0e0f2 · inbound

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt cites this paper.

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt Better speech synthesis through scaling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:39.764367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:39.764367Z digest=sha256:50daf81182c4ec3b21cfb7ed319a65a44d12ac800033c6e39e7ee26c01a5a162

Observation 9951754c-b7b0-4e2a-9983-da0f89676296 · inbound

Semantics-Aware Human Motion Generation from Audio Instructions cites this paper.

Semantics-Aware Human Motion Generation from Audio Instructions Better speech synthesis through scaling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:33.068807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:33.068807Z digest=sha256:a8fd083fc7847acfca0b18d2262145a01396de85a4c35b30bd7a6bf8d4a4ebb1

Observation 40e12eec-8569-4de8-89b0-2d584479c018 · inbound

StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding cites this paper.

StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding Better speech synthesis through scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:33:09.226193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:33:09.226193Z digest=sha256:0841aaeff9b225782aca2a3362bfd466b1fc754355583229d4bef039bd9a081a

Observation efa6ef73-0ed5-4a80-8573-fa8b4f8c2b25 · inbound

Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis cites this paper.

Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis Better speech synthesis through scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:25:26.849481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:25:26.849481Z digest=sha256:3dcfa9d3e7b4392fee39f78baa6e58fd8a296b8f5d66948bbc0bba1bcbf3eb92

Observation b03c3b1c-9fbf-4f63-b925-c5162c052d5f · inbound

De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning Attacks cites this paper.

De-AntiFake: Rethinking the Protective Perturbations Against Voice Cloning Attacks Better speech synthesis through scaling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:19.238058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:29:19.238058Z digest=sha256:63706146727fa75136f34ff8bd8d41ef57173f2b661cf87196ccf4af63bb9098

Observation edcddf70-53f9-47ca-b41a-b5fb3e02e6dd · inbound

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations cites this paper.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Better speech synthesis through scaling

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.512702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.512702Z digest=sha256:210bb5e3431b037a111293a632a2eca066a40ee805353056ef3610362079e336

Observation b830277d-a5b9-48e4-8861-901a9974e525 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report Better speech synthesis through scaling

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:50.966924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:b52a0d7f02e9452f18f0373b24f4dc18c68eb000fac2050cc4eb7f5c2e469d0e

Observation 9f8eb8eb-c408-4fb3-bb91-927714cba40b · inbound

JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 cites this paper.

JWB-DH-V1: Benchmark for Joint Whole-Body Talking Avatar and Speech Generation Version 1 Better speech synthesis through scaling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:28.255179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:28.255179Z digest=sha256:b7b4f7579dfed079ce38b15ae0e38e908933fc765cd94f60aaea50b185cb80c0

Observation 448ecfc7-dea8-42b5-b8e6-c5ee136542be · inbound

TTS-1 Technical Report cites this paper.

TTS-1 Technical Report Better speech synthesis through scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.858858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.858858Z digest=sha256:58267e25bc9154cd4355fbea73eeea02c4c55165c425536efda639ed83ca7289

Observation 42bed4f2-64cc-40a0-803e-8e9e71bf4a9b · inbound

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods cites this paper.

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods Better speech synthesis through scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T12:49:16.507053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:49:16.507053Z digest=sha256:76e6d0d696e90f7f40372bc2d651462b7b1d13c3b101b896b38867145197f8af

Observation 6229081c-1426-49c8-90a8-da36704ba759 · inbound

Next Tokens Denoising for Speech Synthesis cites this paper.

Next Tokens Denoising for Speech Synthesis Better speech synthesis through scaling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:25.341706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:25.341706Z digest=sha256:d2d08c2f0d2e9c4651a79d3422515729c88fd523cbe136a44963792802642fe5

Observation 980f1468-0f6e-48d7-bd95-59f812a4ea20 · inbound

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space cites this paper.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Better speech synthesis through scaling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.044506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.044506Z digest=sha256:1de2c8a11bf00299d2f63603bc46a4f0344b390f82ed815832fea768dcee9b66

Observation f5bf894b-6301-462a-a9bb-ff8ace21d634 · inbound

Large Language Model Data Generation for Enhanced Intent Recognition in German Speech cites this paper.

Large Language Model Data Generation for Enhanced Intent Recognition in German Speech Better speech synthesis through scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T22:52:07.682737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:52:07.682737Z digest=sha256:d2a28abfe6fc2bd9b8fd6dbd6d4761f7c595771ccd73d83f88f088c09bbbf431

Observation 314b1853-2332-4c49-a51b-96a09c774be3 · inbound

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training cites this paper.

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training Better speech synthesis through scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T18:03:40.305722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:03:40.305722Z digest=sha256:b9ee087d48847bbc35087c042347ea1211ed7b368ad99196998377d2aec552ac

Observation d1011716-3d8a-44e4-a33f-e9b35a272e03 · inbound

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan cites this paper.

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan Better speech synthesis through scaling

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:21:00.750276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:40:10.657590Z digest=sha256:3118f87af8590d57d086b8e2f9206d5a8d614b1e251a74203407edb63cd2baee

Observation a759e3db-93c3-416b-b1be-c4cac265d7d1 · inbound

Enhancing Conversational TTS with Cascaded Prompting and ICL-Based Online Reinforcement Learning cites this paper.

Enhancing Conversational TTS with Cascaded Prompting and ICL-Based Online Reinforcement Learning Better speech synthesis through scaling

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:01:01.896689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:50:59.950646Z digest=sha256:4a06054d0a458008164712e5463ca4b98521d20684a479ac7662a6ef7443c35d

Observation d5d29c86-d4cd-4850-963e-10ca7a84a21b · inbound

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning cites this paper.

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning Better speech synthesis through scaling

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:15.990038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T04:35:58.032597Z digest=sha256:527d45acc0e001f882b8a3e320f286f3ebbc31fe4b463ace3c50d1b82a60b084

Observation ecca9889-85c2-48b9-b1e4-0094a650bd7a · inbound

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning cites this paper.

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning Better speech synthesis through scaling

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:46:13.918576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-12T01:43:48.555523Z digest=sha256:80fe6588c953a74118b60c8b26e2d6509bf6686416be915da8ceaacfbe9ac50f

Observation 994559ed-963d-4afc-a376-75312646c751 · inbound

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation cites this paper.

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation Better speech synthesis through scaling

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:35.150832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T15:15:10.060770Z digest=sha256:e2122495cc014842420f1a048fb7fd81c8e019ee4c3a3579fda100ca14bfbaaa

Observation 39f7ca8a-ae34-4a16-aefd-8bbc381cc9bd · inbound

An Evaluation Framework for Text-to-Speech Voice Reconstruction cites this paper.

An Evaluation Framework for Text-to-Speech Voice Reconstruction Better speech synthesis through scaling

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:39:38.696570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T13:16:05.358573Z digest=sha256:2c1874fce7f71177a47818b7347395547d9d6f6a278039254cefcce0497b1897

Observation dfa2c6d2-65a8-4f9f-90d2-750168104c71 · inbound

Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis cites this paper.

Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis Better speech synthesis through scaling

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:10:07.254111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T20:43:29.118351Z digest=sha256:93fd4d08fb7e3289c547c1728950ae795de39dc30382dcdd710486ca3d1ee9a6

Observation d9687248-44c7-475a-9473-379c1881c1da · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Better speech synthesis through scaling

Reference 113

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:55:42.196372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:749639cb781ad454f3731e5b0e677527af37ee72622c0c9d1a3d22f1a3b0adf8