Pith. sign in

Paper Citation Record · LEDGER

SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2406.02328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.02328 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:12:56.625844Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:08:37.504398Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cb511548-a040-4418-975d-3dd8512cc73a · inbound

DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model cites this paper.

DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:56.625844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:56.625844Z digest=sha256:d2b5947406ab32d0d5ecef008c1a102d3a9d2f2ead52c5067aab5eb4bad82331

Observation 3532d1ea-3948-4fe4-a0b4-111f73970903 · inbound

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech cites this paper.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:59.444748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:59.444748Z digest=sha256:85c581bcf430f6a2e2db52e174c23594ad648010f6a8c747d10a2fed5065b726

Observation 9cb08e55-b39c-4f59-ab51-4d28892893d5 · inbound

Adaptive Duration Model for Text Speech Alignment cites this paper.

Adaptive Duration Model for Text Speech Alignment SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:09.476364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:09.476364Z digest=sha256:c36c306b3deaff7512580fe2539c221249feecfe085e68814574b6b92b62ab97

Observation b78d2d07-028f-4631-9aff-1f837a23dda8 · inbound

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking cites this paper.

Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T21:11:51.164939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:11:51.164939Z digest=sha256:7daf59be418b5a55ded7ec2b687c347907592844e1302ba4b80eeac6d298b777

Observation ea9e05b2-2663-4396-bd66-7eaa480045e2 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.175358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:2df49027f21542604006aadd596ce571e04a01483c42a0fbd127b3e4dc0b4f88

Observation 5645eff3-53d7-43e8-8aae-52110c772c1a · inbound

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment cites this paper.

Self-Guidance: Enhancing Neural Codecs via Decoder Manifold Alignment SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T16:08:37.506013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T06:05:26.735340Z digest=sha256:7cd0aeb0ca940ec278223df9e32468f90b8e5403905a16dca2d4bdbd72dcf9c5

Observation 6ca38534-66b4-4f39-af00-cc40940c04a9 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models

Reference 180

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.163389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:162c6377a407e26c418c91a16bae06bef4dcd2271032f1cb23a6b556c451042a