Pith. sign in

Paper Citation Record · LEDGER

DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2406.11427.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.11427 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:37:04.949485Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T12:58:08.367612Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation de99c861-4fa9-4af9-b871-0aa3d7d9c1e7 · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.413952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:a475532941d35b22de7fa390edb92d85473a47ff3021de50f6859ce8dffd4539

Observation 2d0156fb-cd14-41b7-aed1-cbe8a695345c · inbound

Zero-shot Voice Conversion with Diffusion Transformers cites this paper.

Zero-shot Voice Conversion with Diffusion Transformers DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T20:12:09.705643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:12:09.705643Z digest=sha256:c1dade8cf2a9b5ccbc1e4324f03f3b8686b9123b4e2bb45ac7d628fb5354e829

Observation 5fd9797a-1517-42f5-a34d-5fe417ab5401 · inbound

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models cites this paper.

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:19:09.559170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T06:19:09.507440Z digest=sha256:2405559d3db523c26e00dcc8a3c53ecab677fabda3968bd3f43b4ff192942ddc

Observation 8b8859f6-5518-4147-b656-d74c0e71beab · inbound

VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis cites this paper.

VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T00:50:14.346101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:50:14.346101Z digest=sha256:d57779c8b403c430133d4e5f9283b9324424f86973d969c87479ae980230f7cb

Observation 1dca5944-1989-40eb-ac7d-4630ce216ed9 · inbound

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder cites this paper.

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T22:16:17.285213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:16:17.285213Z digest=sha256:407529ba996a6c1f0b205040a4e748ed460d81e18fde452b4c5caff05cceeb04

Observation ddbce4e2-199e-428a-8189-e8123457ed81 · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:25.522200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:7162cb969fd46fac23da36349ee2013cfe5c30492f99607a7ea892ba20cd21e7

Observation 76fed29e-242d-4dfa-b6fc-41b2f8fc9140 · inbound

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech cites this paper.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:55.843335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:55.843335Z digest=sha256:179aa1e197ac8ba7cbfcb978ee47a52271ccdb55c6eab767a66aaacaea9df702

Observation 23782b60-0a99-4477-88f2-c59be7d64832 · inbound

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching cites this paper.

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:50:51.024503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T00:46:39.196042Z digest=sha256:f9a9e877fc867d2b27874edf18889d9f13821e72017b69c4267f2d144b6a57ff

Observation b3eef67e-c8b2-4d44-a2e5-8e5a1da552c8 · inbound

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis cites this paper.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:21.419026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:21.419026Z digest=sha256:e9c4364b90725e30ea85ad60c1678b2071274bb8780d1e010f7345445c074c01

Observation 2c943391-9c5c-40ac-a1ca-27fd64d2e275 · inbound

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration cites this paper.

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T19:11:59.542297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:11:59.542297Z digest=sha256:7bcd417647fe6dc7bee4bdbf40a809d349533802edb06b902ed20aab4f637ff2

Observation 89c0ecd3-ed9d-45c2-9da5-81d777db0f0b · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.183773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:77b2f461877385ee21992317f1b796285dc974fa07deade8bcaee250e61f9407

Observation 0e3aaad8-48ce-4b1e-a5c5-1a1eb2beeeec · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 100

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:f398986af0e4c9b7c1bd657699f41abfd98c9e627999a88384346a1991507ee5

Observation 51c414b0-ffcf-41d2-806b-9a5443fc2255 · inbound

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech cites this paper.

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:33:55.447788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T02:32:26.122526Z digest=sha256:db0fa14cd7937896738bf9979ab59ccf011c1f3e5310a8ebfb4662d891616dd6

Observation b5d83bd4-65ad-4c8b-ac08-24fed19301e4 · inbound

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations cites this paper.

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:58:08.369088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T08:38:51.371500Z digest=sha256:8510a30826c72b8862848e3f7d82e0383fadab438472a64bef08e80457f32b79

Observation 4af0ea1f-98f3-4fe0-9baf-4e6e14e5981a · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 181

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.165963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:7fc815ef99344a744b22c71543d8b929c415eda69b9fccc57cfb13dd2912767d

Observation 42a7f1bd-868f-46a0-872d-4e05a7e68c89 · inbound

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech cites this paper.

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-11T21:24:36.925360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:24:36.925360Z digest=sha256:88e511567066bdad3eb9f3679daf54505ff471c9d8aa6d8000ae99ce64c26933

Observation 0f71aab3-e39d-4989-b3f4-31a948676316 · inbound

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization cites this paper.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.949485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.949485Z digest=sha256:c5178a0e11e7d946c1a827d6a9731011fd06503ad92ec3f1a98c94efab202289