Pith. sign in

Paper Citation Record · LEDGER

BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

As of 3 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2402.08093.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.08093 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T03:50:26.873406Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:17:30.982050Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e2f60dc4-b155-4d77-ad62-d3f221336643 · inbound

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models cites this paper.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.337799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:4541d28fdd8709893ad4638d2094e60c5169d5871ac316eee29d278983c3ea1b

Observation 16b1df01-f4af-463a-8651-d4611665133e · inbound

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models cites this paper.

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:19:09.589168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T06:19:09.507440Z digest=sha256:d5979fd7103edd35556ddd53b342848c055f248d4d46886a8ab14fb4f873347c

Observation 5efae641-450d-46bb-b022-57e8b05441ae · inbound

A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech cites this paper.

A Novel Automatic Framework for Speaker Drift Detection in Synthesized Speech BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:41:00.646707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-10T18:24:54.627215Z digest=sha256:c68b3e511071d1378801a4b564616add3924363bdf107331a299e8e6cb97ccfd

Observation 6085b80e-e5f4-4942-a4ff-fce37384abb4 · inbound

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning cites this paper.

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:41:16.018613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-08T04:35:58.032597Z digest=sha256:3509d8b95d6acb0897c11525c8972302d7cfa18866891d70aa34859fdf5a4894

Observation 60e59582-95e9-4841-88cd-a75413769421 · inbound

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning cites this paper.

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T01:46:13.907857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-12T01:43:48.555523Z digest=sha256:21c30b3f2d8fb87a2d5d22f2adc780f2e7002c0d892881652731461279f97a41

Observation a2492884-b458-442f-8967-03ff1d21b029 · inbound

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis cites this paper.

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:02:43.574934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T18:58:15.288299Z digest=sha256:d878981b495286eb7d284fc9015f1ad2ed1955514a674c0c610c33ca3c658ebd

Observation e5a1bf20-4e87-4599-a4e6-88cd7a38d279 · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.768082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:04d458ff39a02a66876e91ea4fbe8120964190f7c7ea7f9e3a960a477f4f4964

Observation 185f221c-75e4-46f6-ad84-16083576bf82 · inbound

N\"ushuVoice: Reviving the Voice of Endangered N\"ushu with Pitch-Aware Text-to-Speech cites this paper.

N\"ushuVoice: Reviving the Voice of Endangered N\"ushu with Pitch-Aware Text-to-Speech BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:17:30.983342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-27T16:39:49.263497Z digest=sha256:4479842331dcf8ae574396615c47b3a7f0d29a5d3101ce821a80817de5d7d0d3

Observation 67d3e437-f0e2-419f-b999-9e21f8c6829e · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:55:42.217957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:710b178edfecae235d811355c2784a32484204b6529268257a3982282162a495

Observation 06b38db4-bca4-471f-b398-0ccac188d768 · inbound

Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation cites this paper.

Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:55:42.974888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-01T03:32:23.838961Z digest=sha256:130f857f8b8963572982b592c09cb7d019acf670df10305cc9766cc0c1a598e0