Pith. sign in

Paper Citation Record · LEDGER

TTS-1 Technical Report

As of 7 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 4 inbound Pith citation observations for arXiv:2507.21138.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21138 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:02:03.077732Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:36:31.523597Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:19:02.942476Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1116d9ff-1bb6-45fe-8c3b-2d690506a98d · outbound

This paper cites The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage.

TTS-1 Technical Report The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.830714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.830714Z digest=sha256:7952e7bcc595ef9281ebd48ef90772eee0d064351621dd8bff9a07e4f7566c75

Observation b6165c63-f5b3-490c-908c-7f7aa6a4e881 · outbound

This paper cites Yodas: Youtube-oriented dataset for audio and speech.

TTS-1 Technical Report Yodas: Youtube-oriented dataset for audio and speech

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:04.059235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.837634Z digest=sha256:58d12f0dad4516b06ac73c04fb9a0916b9d3d875f94a58f7ba581edf98845a52

Observation 9a8e03bc-ef3b-494a-8385-5600549cb4dc · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

TTS-1 Technical Report Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:04.040769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.843161Z digest=sha256:35791ee39fc9fdb1aec20c265b8a7eda8e344bba52b663572db9c6b7e858f19f

Observation a5392520-ba95-430c-a8b1-d442e5564146 · outbound

This paper cites Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models.

TTS-1 Technical Report Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:04.025179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.847985Z digest=sha256:e5c825bd923759c625e7ba4a5dbe33db4c811c2857f1c8c84d741cfcfa0578cf

Observation d993e778-bc6a-4974-a47b-8ab4e75e5f8b · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech.

TTS-1 Technical Report Fastspeech: Fast, robust and controllable text to speech

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:04.008207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.854081Z digest=sha256:66c9d26dac0da5a03e2688e5e66fbdbf93a1df5d493d16dcde2e510f12f478e3

Observation 448ecfc7-dea8-42b5-b8e6-c5ee136542be · outbound

This paper cites Better speech synthesis through scaling.

TTS-1 Technical Report Better speech synthesis through scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.858858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.858858Z digest=sha256:96bdf7e1f4f8a51cf683935ac0f02e6a4f391e21aea9b0e3432ee91f1b446a4c

Observation e9573b28-862d-43f8-928b-4878a58d2629 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech.

TTS-1 Technical Report Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.988145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.865173Z digest=sha256:20f61666d015e95d33fb34944750d42f4fe12682b09f43a1509777e338f14eed

Observation e204a9e3-3816-48fe-84cf-4e4f73ff3a0c · outbound

This paper cites V oicebox: Text-guided multilingual universal speech generation at scale.

TTS-1 Technical Report V oicebox: Text-guided multilingual universal speech generation at scale

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.965541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.870328Z digest=sha256:f1e69ff6585468e6f473133be53f0d31f24e67de77d9ed88dd23b9bce8ae5405

Observation 3fb837f9-85f5-43b5-add2-e03a99597d36 · outbound

This paper cites Language models are unsupervised multitask learners.

TTS-1 Technical Report Language models are unsupervised multitask learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.875785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.875785Z digest=sha256:4e523e0f9317ba507c0b5555c0a8b7f5ec6ef3f71a2612015b338b0e18b8c2b4

Observation 7c2d9832-c784-4272-b9a4-311689e7b0d2 · outbound

This paper cites Training Compute-Optimal Large Language Models.

TTS-1 Technical Report Training Compute-Optimal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.880909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.880909Z digest=sha256:19d10d6159aa13ff23a8da7fea8dbb2231b696a6854b5f6dab9bd06a00ca93b1

Observation d96c6f96-d5ae-4aa6-b75b-63251abf5781 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

TTS-1 Technical Report Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.886470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.886470Z digest=sha256:e722ddd35223fdea5de0db0c4d1ffb81efdbfb8dc67b0cefede81b105c92ef97

Observation bc5e8aef-029b-4609-8889-5fef279eb141 · outbound

This paper cites MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder.

TTS-1 Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.891456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.891456Z digest=sha256:6c1804a5c769a2296ff9cdd549e2877afea2d3cb729b3478bd9949cbaa5b2047

Observation aed9d272-a437-42ff-bec8-c77046dd3603 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

TTS-1 Technical Report Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.897698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.897698Z digest=sha256:a5e36f2b5c770dc27ca85b8f175dbb0796b0bf412b2f4cd81565c717dc05b767

Observation a3c777f1-cd98-4493-aae4-b99d5ea637dd · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

TTS-1 Technical Report CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.904024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.904024Z digest=sha256:4963e0cbcb7fb3c2c932c229ba172508c8b124379011f9cc1d9e778f5f4a70e4

Observation a2cd3a78-d85f-44a1-9c4b-f1c80281eac1 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

TTS-1 Technical Report CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.908687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.908687Z digest=sha256:020d5ddc55cf565b5d7313ced7ed80b43303c56eb421b7f9b450040694991f88

Observation bd31abf3-cdc3-4957-9bde-e4ca9251251f · outbound

This paper cites Redpajama: an open dataset for training large language models.

TTS-1 Technical Report Redpajama: an open dataset for training large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.926487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.913946Z digest=sha256:93e37fcf9fcfc51753fe9adce45bfa7e58725f705eb74ec4ce6d5b13d7c97852

Observation 669d28f8-b554-4228-a0ee-fdb73850def1 · outbound

This paper cites Open instruction generalist (oig) dataset.

TTS-1 Technical Report Open instruction generalist (oig) dataset

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.903451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.919362Z digest=sha256:ce94fb166c496ae738cecf16885b9a035aa05cdc82daf3b4ccd63b198d4a813f

Observation 1f6b538f-24df-4cef-b876-74eb4f3f71dc · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TTS-1 Technical Report DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.924574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.924574Z digest=sha256:a57e7012dd0cc922efad418064311334f03d7dafc14d549e7c1b24c8e1e62068

Observation 5a853971-6149-46e6-9863-4098efe1267e · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

TTS-1 Technical Report Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.930108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.930108Z digest=sha256:cb33c2901c6f46b7f573efefca96677faad752d84fe59a6c4eadc9db7e051a09

Observation 61e20d6f-5552-4c0c-a9a8-b4a2d805e76e · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing.

TTS-1 Technical Report Wavlm: Large-scale self-supervised pre-training for full stack speech processing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.884270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.935030Z digest=sha256:dd7f42c9b0b4297818084871e43388dbde6ed8bf9776369972fb28e1f1673168

Observation befd186d-6c49-493d-8b30-f788cea54faa · outbound

This paper cites Dnsmos p.

TTS-1 Technical Report Dnsmos p

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.865444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.940134Z digest=sha256:5b3285ce2d52751c3ec28cc652810c3aea2cf850185364ddd640c73139e8a2d5

Observation 3d900436-35b3-4f17-b713-eba6b6f78fa3 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

TTS-1 Technical Report Lora: Low-rank adaptation of large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.944597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.944597Z digest=sha256:36f7fc458ea231fcb09d6218b9add01878e283424aa92e7eb141026af937b19e

Observation bdf5b87d-9598-4cab-88a2-5c88c38d010a · outbound

This paper cites Deep residual learning for image recognition.

TTS-1 Technical Report Deep residual learning for image recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.949788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.949788Z digest=sha256:a676074fd1326b39962c3224e08ef81a60e9a27107ea39849d12827f82cf2d28

Observation 089d6aab-8a5f-4074-bdbd-071b1dd9a10c · outbound

This paper cites The Llama 3 Herd of Models.

TTS-1 Technical Report The Llama 3 Herd of Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.954136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.954136Z digest=sha256:90bab65e90c327391cadc726d7c4a5e0bdf3f723ce005811251a44da3c2d8fdf

Observation 6e452096-9f3d-4855-8a96-101daea28de3 · outbound

This paper cites Matrix multiplication background user’s guide.

TTS-1 Technical Report Matrix multiplication background user’s guide

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.825369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.959872Z digest=sha256:5924a0d7cb1dcd6814682789cb83a7476a03daee96fc5a8e843fee5272ebf66d

Observation 63165417-09a3-4464-b43e-d01040b7e49f · outbound

This paper cites Initializing new word embeddings for pretrained language models.

TTS-1 Technical Report Initializing new word embeddings for pretrained language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.808174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.964733Z digest=sha256:0ff1ed5e9cee1ad935826923d3378a5e888b3d8e68869199cce3c820958232e5

Observation dee61c16-a693-413d-baec-a8d71bc82094 · outbound

This paper cites BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec.

TTS-1 Technical Report BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.970136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.970136Z digest=sha256:0afc670f9b38b20a993324baa02c9475c6de777cccb01822f1af24ab1a9ac15d

Observation e7372153-22da-4005-9a06-a28caf869736 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

TTS-1 Technical Report Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.788491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.974412Z digest=sha256:4704a241de8885f68c8c8001afeaed4174779693af5ed63fe553fe3b6af12b71

Observation 331a3e81-ee3b-4418-9bf2-aac95671b668 · outbound

This paper cites High Fidelity Neural Audio Compression.

TTS-1 Technical Report High Fidelity Neural Audio Compression

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.980832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.980832Z digest=sha256:0ca2b0e1c5db43a5833377ab100ca4df666c89eb0d3825a34c674dd8f74ad3dd

Observation dfac856f-3877-4bf0-9bdb-d925ad897b44 · outbound

This paper cites Adding Instructions during Pretraining: Effective Way of Controlling Toxicity in Language Models.

TTS-1 Technical Report Adding Instructions during Pretraining: Effective Way of Controlling Toxicity in Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.985709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.985709Z digest=sha256:fe9a0fbe73118c9f1d15a87767eec9adcf8bb4cd9724a01154536c896663049b

Observation b2d0e952-b052-4dc2-8a25-3c9367e1cfe1 · outbound

This paper cites A neural probabilistic language model.

TTS-1 Technical Report A neural probabilistic language model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.990541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.990541Z digest=sha256:7a552087492b76912db66955df62cf8b39013b20bf614c59275cbd40baf574bd

Observation 9133858a-f8b9-420b-8e82-19916efa00be · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

TTS-1 Technical Report FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.997072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.997072Z digest=sha256:337ff214d515efea43f47d711300e2f187130ce8cb883a408a0d4011b779ad85

Observation 29c6c949-fa70-4f7c-8811-4deab52cb8f1 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

TTS-1 Technical Report Adam: A Method for Stochastic Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.001891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.001891Z digest=sha256:40f0b25c844704f53e46c0516dd2d3ecd41ceed48440598579236bf82e93b061

Observation dc53bb10-0b74-41a7-93a0-bac436025d36 · outbound

This paper cites PyTorch Distributed: Experiences on Accelerating Data Parallel Training.

TTS-1 Technical Report PyTorch Distributed: Experiences on Accelerating Data Parallel Training

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.006360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.006360Z digest=sha256:79c9de75bcf20797f33f0164b79363c10184eb34ae927bc2d413511529908a9c

Observation d962fbbe-3828-4961-9347-cf0e52fb0db0 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

TTS-1 Technical Report PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.011140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.011140Z digest=sha256:68863d5a348877fc6e73ee44c41a6b31f7507a8b30ebf6c660022a54912640b0

Observation e27dba3b-4a19-4c67-a1b4-2ed803652cb5 · outbound

This paper cites Parler-tts.

TTS-1 Technical Report Parler-tts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.755019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:03.016065Z digest=sha256:0c02eeb8d21c16981f8ef400a29100225da9d071a1b957721074c17ac1905bfa

Observation 8a5c05da-c5d0-4796-bf6d-9303960c483b · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

TTS-1 Technical Report Zero: Memory optimizations toward training trillion parameter models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.737535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:03.021572Z digest=sha256:64297bd439a94d4748278022bb8a6fa084a14d608631e0aad0462dd6d16ba38d

Observation 52ded5a3-2265-49cd-b401-61758dbf0392 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

TTS-1 Technical Report CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.026897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.026897Z digest=sha256:42b36cd8a9b1e4c465b7effc5c4aefdf08f11428247ad26365215da4ee7e5fbe

Observation 3640e72a-bbf6-47a0-89f5-b238dca3c766 · outbound

This paper cites Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback.

TTS-1 Technical Report Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.032073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.032073Z digest=sha256:e76746db0c18ea8061d4ba6300ccb3bc513793b66356e9f68daf3cb6e386faea

Observation 4a88a8cd-6997-471a-8f21-bc960cf587e3 · outbound

This paper cites DeepSeek-V3 Technical Report.

TTS-1 Technical Report DeepSeek-V3 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.037268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.037268Z digest=sha256:b7481e384b327aebf694ef8f489b32c6945a49d08012d8a3277bf4b091d0480f

Observation 63c76d66-c01b-4edd-9a8a-2caebbdb197f · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

TTS-1 Technical Report DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.042360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.042360Z digest=sha256:ceb8fcc27c9006379dd77cf62cb18e9b13e02f2c1c2f0a624e51317b9fc47dc7

Observation 4ca58588-1f1f-4943-8194-60f8dbc61c15 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

TTS-1 Technical Report Direct preference optimization: Your language model is secretly a reward model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.047248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.047248Z digest=sha256:be2521821abccab1ecefd0719fd5825060fe2d7756384ad84af932c55fea1ee4

Observation f720f069-b85f-48e8-afe8-ae422da1397e · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

TTS-1 Technical Report Understanding R1-Zero-Like Training: A Critical Perspective

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.052195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.052195Z digest=sha256:d2c23d2239126a359d88ce24fbe9e59be9eaa85497212781f98eaf1512a6fda0

Observation b9b904ce-8b94-46ad-9c3d-99590258e76e · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

TTS-1 Technical Report Robust speech recognition via large-scale weak supervision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.057567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.057567Z digest=sha256:9b647ceef7b24d31c3f33f4f5a7ca4d9f03ad1a880c4801e0661029267134a74

Observation 6d2518ba-0c5b-4908-ace2-7d94f7171699 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

TTS-1 Technical Report HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.063141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.063141Z digest=sha256:fffa2e6a092c59591549bf81acfd0f264cc3506d9f84d7b1b1b59aa4b626ef1d

Observation bd80dfc2-0d2c-400a-818f-ebd963c3eeab · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

TTS-1 Technical Report PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.068483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.068483Z digest=sha256:562ec43d0510f79d1de25260599007598817210c9725a60e853d25dfe944321b

Observation a606eda4-cfd8-44ee-85bf-7b99585a2b9b · outbound

This paper cites Pytorch lightning.

TTS-1 Technical Report Pytorch lightning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.690672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:03.073406Z digest=sha256:b1da5b934a23071e68379aaa2a71c7ade41c02f2c48d24e802f104004da34e34

Observation c88b760d-9f09-4e4e-a31c-fb9f6ae7cc7b · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

TTS-1 Technical Report Efficient memory management for large language model serving with pagedattention

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.674403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:03.077732Z digest=sha256:7c89ec3559d22ea5483877d2ad07d59cf904a17d9a1d1d185dd3fedda997743b

Pith citing papers

Observation 64fed19a-4d58-49a6-867f-8977988efecd · inbound

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability cites this paper.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability TTS-1 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:31.523597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:31.523597Z digest=sha256:e7407946e83e90b39e92f262ce9655e0cbc1bb4648140557bc94e4d26dc4f6fe

Observation b76085d3-b265-4c38-beb9-aa07f37b5812 · inbound

JaiTTS: A Thai Voice Cloning Model cites this paper.

JaiTTS: A Thai Voice Cloning Model TTS-1 Technical Report

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:28.850669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-07T06:34:45.357695Z digest=sha256:6ce68ffddbb3fe5cae00b67f4465bda7875c99b00eb5e3d220bc520b75cc1f91

Observation cca52b6c-b63b-414c-9517-cb5c8bc46c2a · inbound

JaiTTS: A Thai Voice Cloning Model cites this paper.

JaiTTS: A Thai Voice Cloning Model TTS-1 Technical Report

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:16:26.706638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T03:09:24.541657Z digest=sha256:be08a18e2375077612a4c7a829eb13f6a3cb0c968a74029263b37c12f54defb8

Observation cc3f5c45-0029-4f23-a0c6-f49bfa10d585 · inbound

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs cites this paper.

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs TTS-1 Technical Report

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.945574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T22:46:47.786929Z digest=sha256:d3db8e3831fa0be52fd1b89688bdb980fe9d631235ec9dffa5e742287ee488c1