Pith. sign in

Paper Citation Record · LEDGER

TTS-1 Technical Report

As of 18 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 4 inbound Pith citation observations for arXiv:2507.21138.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21138 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:02:03.077732Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:36:31.523597Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:19:02.942476Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1116d9ff-1bb6-45fe-8c3b-2d690506a98d · outbound

This paper cites The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage.

TTS-1 Technical Report The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.830714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.830714Z digest=sha256:734c79d3c27e3d7fd0d8f667bbd46b081e4dba94a80b676a1f50aaf3bcb15687

Observation b6165c63-f5b3-490c-908c-7f7aa6a4e881 · outbound

This paper cites Yodas: Youtube-oriented dataset for audio and speech.

TTS-1 Technical Report Yodas: Youtube-oriented dataset for audio and speech

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:04.059235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:02:02.837634Z digest=sha256:84725f8db698d516a6491e0d622f7119c4a2f24f91f7867d53dad31e572292cc

Observation 9a8e03bc-ef3b-494a-8385-5600549cb4dc · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

TTS-1 Technical Report Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:04.040769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:02:02.843161Z digest=sha256:c6bdc70fa77c108d951ce43ea23bf358eeab5e0c9d2f707b2919cab969f66750

Observation a5392520-ba95-430c-a8b1-d442e5564146 · outbound

This paper cites Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models.

TTS-1 Technical Report Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:04.025179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:02:02.847985Z digest=sha256:57e68222b1944c7d74a022590852c7cbcb0974fc3917b740dfadac1146bc26ee

Observation d993e778-bc6a-4974-a47b-8ab4e75e5f8b · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech.

TTS-1 Technical Report Fastspeech: Fast, robust and controllable text to speech

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:04.008207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:02:02.854081Z digest=sha256:35b2ba8966ad7fd5f342aa2f8d052033e3eca5176a74729b40d594954292836e

Observation 448ecfc7-dea8-42b5-b8e6-c5ee136542be · outbound

This paper cites Better speech synthesis through scaling.

TTS-1 Technical Report Better speech synthesis through scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.858858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.858858Z digest=sha256:bbdf0fc19041796d6cc0ecf74f085e209c63af6c7a8a44868132cd36f3844d46

Observation e9573b28-862d-43f8-928b-4878a58d2629 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech.

TTS-1 Technical Report Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.988145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:02:02.865173Z digest=sha256:0821d17bd6f5ef913df9d6753f95421c0f633db2797b9c27af9c7e768cbe7257

Observation e204a9e3-3816-48fe-84cf-4e4f73ff3a0c · outbound

This paper cites V oicebox: Text-guided multilingual universal speech generation at scale.

TTS-1 Technical Report V oicebox: Text-guided multilingual universal speech generation at scale

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.965541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:02:02.870328Z digest=sha256:792e97d1cdfc881343bea28c9cde75b60647d1ddf957d828562964993831a84e

Observation 3fb837f9-85f5-43b5-add2-e03a99597d36 · outbound

This paper cites Language models are unsupervised multitask learners.

TTS-1 Technical Report Language models are unsupervised multitask learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.875785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.875785Z digest=sha256:96b18fe1d429439a4f8deaae94fa63e44ca126e61d2e61d90f402cffa423fc28

Observation 7c2d9832-c784-4272-b9a4-311689e7b0d2 · outbound

This paper cites Training Compute-Optimal Large Language Models.

TTS-1 Technical Report Training Compute-Optimal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.880909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.880909Z digest=sha256:646c6c61ca450597dd32244599ddfec1511d7738b00272bdcd5ca995b758bad6

Observation d96c6f96-d5ae-4aa6-b75b-63251abf5781 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

TTS-1 Technical Report Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.886470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.886470Z digest=sha256:374cb3560ca879b621a32eaa7bdd651d2bbfd9e7075566d72f65ad17658e77b1

Observation bc5e8aef-029b-4609-8889-5fef279eb141 · outbound

This paper cites MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder.

TTS-1 Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.891456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.891456Z digest=sha256:704714ec0f379dfc0bcca3e7077d416de9530396543c79f33e46d0f8a1517220

Observation aed9d272-a437-42ff-bec8-c77046dd3603 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

TTS-1 Technical Report Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.897698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.897698Z digest=sha256:6f5b352224ede4ea405556b3eab42cfd09560d0930f4420f38473efce8f5ee31

Observation a3c777f1-cd98-4493-aae4-b99d5ea637dd · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

TTS-1 Technical Report CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.904024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.904024Z digest=sha256:23c7ddb7e2c495e69c7cb54eb3d62e34e2165dbc008df969a25d3b7a9436b5bd

Observation a2cd3a78-d85f-44a1-9c4b-f1c80281eac1 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

TTS-1 Technical Report CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.908687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.908687Z digest=sha256:c2c4b17ddff551df8b35aa943813c0f9c60406d5351b70e6065987057ea8393c

Observation bd31abf3-cdc3-4957-9bde-e4ca9251251f · outbound

This paper cites Redpajama: an open dataset for training large language models.

TTS-1 Technical Report Redpajama: an open dataset for training large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.926487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:02:02.913946Z digest=sha256:85a5fcb4bc4da491d94fd64ec43d58da3d517378ad49d5de941f91c65fd03d32

Observation 669d28f8-b554-4228-a0ee-fdb73850def1 · outbound

This paper cites Open instruction generalist (oig) dataset.

TTS-1 Technical Report Open instruction generalist (oig) dataset

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.903451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:02:02.919362Z digest=sha256:f510c4ef19d5bf3da72298bb0f9d17d042c3055daf889accfceaab351e5440f6

Observation 1f6b538f-24df-4cef-b876-74eb4f3f71dc · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TTS-1 Technical Report DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.924574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.924574Z digest=sha256:3d13023e2a33f24607d5488ebe6d07580b6f5e1269d26a54892e87f1b396f44f

Observation 5a853971-6149-46e6-9863-4098efe1267e · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

TTS-1 Technical Report Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.930108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.930108Z digest=sha256:69c90959ce13ccba5947883ee7baafa65f9d3ea9bb4e1fdce5706ca43dfd3f68

Observation 61e20d6f-5552-4c0c-a9a8-b4a2d805e76e · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing.

TTS-1 Technical Report Wavlm: Large-scale self-supervised pre-training for full stack speech processing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.884270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:02:02.935030Z digest=sha256:dd1a8a5ee315e2257b9c6e601d4a5e318cbd6942e760991937c470068c538000

Observation befd186d-6c49-493d-8b30-f788cea54faa · outbound

This paper cites Dnsmos p.

TTS-1 Technical Report Dnsmos p

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.865444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:02:02.940134Z digest=sha256:e7d554e2d8db63f4feb6323cfbfae8045928a2e444761ff6b8bd576d3b024820

Observation 3d900436-35b3-4f17-b713-eba6b6f78fa3 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

TTS-1 Technical Report Lora: Low-rank adaptation of large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.944597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.944597Z digest=sha256:ab2ae1291f185c7ef7f8e5bb9415dd1f8a9166a888b6e3110cdc0026c4d215d9

Observation bdf5b87d-9598-4cab-88a2-5c88c38d010a · outbound

This paper cites Deep residual learning for image recognition.

TTS-1 Technical Report Deep residual learning for image recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.949788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.949788Z digest=sha256:189aa12c639aa6d7803a093fa8f21dc1e7e07dd8959cf8235dff990d0c69416a

Observation 089d6aab-8a5f-4074-bdbd-071b1dd9a10c · outbound

This paper cites The Llama 3 Herd of Models.

TTS-1 Technical Report The Llama 3 Herd of Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.954136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.954136Z digest=sha256:b71e10d7a9a2dfc451b51fe2c192628fdd8cb9c67fb94a81144bf1f035735696

Observation 6e452096-9f3d-4855-8a96-101daea28de3 · outbound

This paper cites Matrix multiplication background user’s guide.

TTS-1 Technical Report Matrix multiplication background user’s guide

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.825369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:02:02.959872Z digest=sha256:700da88b4094ba67bafdc19f1f5af37f9b63d13a1067e2ac6361d83bd2fe60e5

Observation 63165417-09a3-4464-b43e-d01040b7e49f · outbound

This paper cites Initializing new word embeddings for pretrained language models.

TTS-1 Technical Report Initializing new word embeddings for pretrained language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.808174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:02:02.964733Z digest=sha256:14a8f91999783bb5e89317ecea5087c4ed7140e20013368e660aded24e0f995f

Observation dee61c16-a693-413d-baec-a8d71bc82094 · outbound

This paper cites BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec.

TTS-1 Technical Report BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.970136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.970136Z digest=sha256:499b81f188d209efc8942f3e805077b527e67b74e4421eed381f4c36e58a6460

Observation e7372153-22da-4005-9a06-a28caf869736 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

TTS-1 Technical Report Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.788491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:02:02.974412Z digest=sha256:6a9a393efe4e4ffb81eb35ee3a732c65cfef96fb4bd98fdcaf987de0d6659224

Observation 331a3e81-ee3b-4418-9bf2-aac95671b668 · outbound

This paper cites High Fidelity Neural Audio Compression.

TTS-1 Technical Report High Fidelity Neural Audio Compression

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.980832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.980832Z digest=sha256:45943897beca54e636920b21384b0b34bd1c049c6d0c6d5a190f4a0ea864aeb5

Observation dfac856f-3877-4bf0-9bdb-d925ad897b44 · outbound

This paper cites Adding Instructions during Pretraining: Effective Way of Controlling Toxicity in Language Models.

TTS-1 Technical Report Adding Instructions during Pretraining: Effective Way of Controlling Toxicity in Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.985709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.985709Z digest=sha256:13cd85e96dfa1ed86a4f9c0775a0616dcca41ce8b2c6a55333e94a1c56809bba

Observation b2d0e952-b052-4dc2-8a25-3c9367e1cfe1 · outbound

This paper cites A neural probabilistic language model.

TTS-1 Technical Report A neural probabilistic language model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.990541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.990541Z digest=sha256:6b0a1139fe9a1f6c4afd12dbc11b4560fbc5e88995d96e8c650b9ac02264bf61

Observation 9133858a-f8b9-420b-8e82-19916efa00be · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

TTS-1 Technical Report FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.997072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.997072Z digest=sha256:12a26bbe8d549deb34f9fba011662d70977fa414f53bccae09d7d2285459279a

Observation 29c6c949-fa70-4f7c-8811-4deab52cb8f1 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

TTS-1 Technical Report Adam: A Method for Stochastic Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.001891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.001891Z digest=sha256:6cd97cba66574d87078d33a15aeef767ee26b2c40859f8efc397f2020b158439

Observation dc53bb10-0b74-41a7-93a0-bac436025d36 · outbound

This paper cites PyTorch Distributed: Experiences on Accelerating Data Parallel Training.

TTS-1 Technical Report PyTorch Distributed: Experiences on Accelerating Data Parallel Training

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.006360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.006360Z digest=sha256:794b50314e137211904bdb70dd4910f0b5bb2ade07aaad52ff1f8f4ef001506f

Observation d962fbbe-3828-4961-9347-cf0e52fb0db0 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

TTS-1 Technical Report PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.011140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.011140Z digest=sha256:5a6690a2e24e6e5a4a88905051de438c242f91c2f0f8def45e7e3754d13e68e5

Observation e27dba3b-4a19-4c67-a1b4-2ed803652cb5 · outbound

This paper cites Parler-tts.

TTS-1 Technical Report Parler-tts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.755019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:02:03.016065Z digest=sha256:22122b94214363eb832f60113f7845d7d4a4cbed6a08c851f753831f561a490d

Observation 8a5c05da-c5d0-4796-bf6d-9303960c483b · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

TTS-1 Technical Report Zero: Memory optimizations toward training trillion parameter models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.737535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:02:03.021572Z digest=sha256:4ff483d23e952b53aa117340b36622ef99149957d42cdb2bdfd9f15273449d4e

Observation 52ded5a3-2265-49cd-b401-61758dbf0392 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

TTS-1 Technical Report CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.026897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.026897Z digest=sha256:4c56fb55c8b597517f115c3b63f85c0d158f6178f02601d5a835ce257035233b

Observation 3640e72a-bbf6-47a0-89f5-b238dca3c766 · outbound

This paper cites Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback.

TTS-1 Technical Report Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.032073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.032073Z digest=sha256:5f3e4a1808e7b225d08bdf9c11c78c3d4903912aa77e3f543c30f1a25ee01475

Observation 4a88a8cd-6997-471a-8f21-bc960cf587e3 · outbound

This paper cites DeepSeek-V3 Technical Report.

TTS-1 Technical Report DeepSeek-V3 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.037268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.037268Z digest=sha256:cea29d50f76448a9da41fb05fce7eae9cfdc639afb60e9115d27db57b0775409

Observation 63c76d66-c01b-4edd-9a8a-2caebbdb197f · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

TTS-1 Technical Report DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.042360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.042360Z digest=sha256:a8268376b7e28d08e7566c2f6763be797d26dfae03e8f4aecc906da24366d0bd

Observation 4ca58588-1f1f-4943-8194-60f8dbc61c15 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

TTS-1 Technical Report Direct preference optimization: Your language model is secretly a reward model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.047248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.047248Z digest=sha256:5512494d6d313ae811e438cbf439b90c3e73ac69195ad382e614aff3691fa6b1

Observation f720f069-b85f-48e8-afe8-ae422da1397e · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

TTS-1 Technical Report Understanding R1-Zero-Like Training: A Critical Perspective

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.052195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.052195Z digest=sha256:26110101ec69a3ddcf4b5d32307b83d623a98b92a15f329b774014afe5c5155f

Observation b9b904ce-8b94-46ad-9c3d-99590258e76e · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

TTS-1 Technical Report Robust speech recognition via large-scale weak supervision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.057567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.057567Z digest=sha256:a2e04d5d671dfe89efccc885909c0c282c9358a7fefd8bdb36c94117a8736aa0

Observation 6d2518ba-0c5b-4908-ace2-7d94f7171699 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

TTS-1 Technical Report HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.063141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.063141Z digest=sha256:09f310aacdba099e339b94b7992ca091024d89ddc834d930b52ea21a5819bc30

Observation bd80dfc2-0d2c-400a-818f-ebd963c3eeab · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

TTS-1 Technical Report PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.068483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.068483Z digest=sha256:fc50d7db02d9b8e592c6463f76a2a993be2c4e6c9251b4b51cd8d556d0ec75f4

Observation a606eda4-cfd8-44ee-85bf-7b99585a2b9b · outbound

This paper cites Pytorch lightning.

TTS-1 Technical Report Pytorch lightning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.690672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:02:03.073406Z digest=sha256:61bf3f9793b8464d7dbf2eb4ab2f9433b3d0a38e882f38cd5cac0da5b95889e6

Observation c88b760d-9f09-4e4e-a31c-fb9f6ae7cc7b · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

TTS-1 Technical Report Efficient memory management for large language model serving with pagedattention

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.674403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T15:02:03.077732Z digest=sha256:bb849ba3b1e0f362e37e029d748fd409050e3acafff524ec49c3333332e6d7fc

Pith citing papers

Observation 64fed19a-4d58-49a6-867f-8977988efecd · inbound

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability cites this paper.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability TTS-1 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:31.523597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:31.523597Z digest=sha256:2474b2ab69bf0140efe544a13ea122474cbf7382225815bc79a69de4bf173382

Observation b76085d3-b265-4c38-beb9-aa07f37b5812 · inbound

JaiTTS: A Thai Voice Cloning Model cites this paper.

JaiTTS: A Thai Voice Cloning Model TTS-1 Technical Report

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:28.850669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-07T06:34:45.357695Z digest=sha256:cfe3e67ea879dbe91404e8c00e7704227cc03e0ab9eafd42b610e50f8b68d924

Observation cca52b6c-b63b-414c-9517-cb5c8bc46c2a · inbound

JaiTTS: A Thai Voice Cloning Model cites this paper.

JaiTTS: A Thai Voice Cloning Model TTS-1 Technical Report

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:16:26.706638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-08T03:09:24.541657Z digest=sha256:78e6d3158cfef8b4218a970c9465d89ab0e3a7201d570587546c15bec28ae3a2

Observation cc3f5c45-0029-4f23-a0c6-f49bfa10d585 · inbound

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs cites this paper.

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs TTS-1 Technical Report

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.945574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T22:46:47.786929Z digest=sha256:fa1b9f44388efda601028cd944969c60d6aafe27107f576ee9b3286a1973405d