Pith. sign in

Paper Citation Record · LEDGER

dots.tts Technical Report

As of 7 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 1 inbound Pith citation observation for arXiv:2606.07080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.07080 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T21:10:25.911203Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T08:35:51.971819Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact18
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a227bf2f-6592-466b-9663-aec08fa98698 · outbound

This paper cites OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models.

dots.tts Technical Report OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T19:57:20.086928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:bc0a6d7f628d2051c3afa3d2ef2cb22974b89a5345328b04ee358c9d003e04e2

Observation d8150420-e9be-4467-9b82-fef24412903f · outbound

This paper cites Longcat-audiodit: High-fidelity diffusion text-to-speech in the waveform latent space.

dots.tts Technical Report Longcat-audiodit: High-fidelity diffusion text-to-speech in the waveform latent space

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.084190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:b97c41c847ad256927b228aa57f565613868a2b5e68809b5306301fbf2f850d9

Observation f77e995f-57c2-4b82-baba-14104bf02e86 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

dots.tts Technical Report CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T19:57:20.078690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:4206c28026346bfa2041be5b3dbdda80092a55f5e924dba5b3c8a4b2f24fed33

Observation 225c645c-38b0-424d-9aa8-3d8c79502e02 · outbound

This paper cites Qwen3-TTS Technical Report.

dots.tts Technical Report Qwen3-TTS Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:57:20.067543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:3e546262bc0b673b9b294ebe2f44dd1f6456925fc268bfda5c038e201a161dac

Observation acec61fb-0e67-4519-b0db-1c0c0d69e4f3 · outbound

This paper cites VibeVoice Technical Report.

dots.tts Technical Report VibeVoice Technical Report

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.073533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:eb986c3e38b4e48f8c4f1c4858ee5904245abc5bed4ba32ec7e69ca9243160e8

Observation 3209b1e8-4ec5-4a6a-975c-a756751fd2ed · outbound

This paper cites V oxcpm: Tokenizer-free tts for context-aware speech generation and true-to-life voice cloning.

dots.tts Technical Report V oxcpm: Tokenizer-free tts for context-aware speech generation and true-to-life voice cloning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.057816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:495b82a3e7486a1908c0a454e9c11336d0239acfc520d64ded2ee8043b7eeb5d

Observation efba0c9f-6502-4976-9253-af1178854f41 · outbound

This paper cites Autoregressive Diffusion Transformer for Text-to-Speech Synthesis.

dots.tts Technical Report Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.039783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:5e6708af536304b283aa478ef92f97e5f16275b949983e41bc89fcdc6b2b2f9e

Observation 34453ff6-37d0-474a-ba2a-82c01b6763b0 · outbound

This paper cites Borderless Long Speech Synthesis.

dots.tts Technical Report Borderless Long Speech Synthesis

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:57:20.089737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:cd9fa4487ea286b7cfde6dfe6d476564d22260f6e277fe0a46bc1c8af82937b1

Observation 0fa92e1a-041f-4bd0-a519-4c4aa1fe2098 · outbound

This paper cites HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding.

dots.tts Technical Report HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:57:20.060177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:df4c2de71c3f32303bff5ecfe296f0ef2aae48716d449622a0a2ee9bc5034c61

Observation a4cabc65-01f3-4afe-a009-3bd74711a548 · outbound

This paper cites SOAR: Self-Correction for Optimal Alignment and Refinement in Diffusion Models.

dots.tts Technical Report SOAR: Self-Correction for Optimal Alignment and Refinement in Diffusion Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:57:20.063637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:fc34d5a7047df5f195610449e1e71b1192116b76643f4c7a44464b90840a93be

Observation 5826c81b-871c-4474-8f24-be0c20aeceb9 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

dots.tts Technical Report Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T19:57:20.079050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:8b94f22aa4a9dd4e1f5e3eb35cbbef1c36a6484650bdb645b18eaf29f776b6ec

Observation d2ae9d75-0bc9-4f8c-b763-dce5fdfe8bc9 · outbound

This paper cites Qwen2.5 Technical Report.

dots.tts Technical Report Qwen2.5 Technical Report

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:57:20.054732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:fca9610ebf7713957838d8e25925d507a856902e112e832c13df968c079e7674

Observation dd263aa1-7927-49e5-9d92-a2a16a60a19e · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

dots.tts Technical Report CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:57:20.062535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:cc325524aaf3c2b69a9ca0c021e1c878706f829b04c8b43764c9f194159f366c

Observation 4af40350-8541-4072-9114-7128ce4677d3 · outbound

This paper cites FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot.

dots.tts Technical Report FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.057397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:23ef0864ff06ba7901d19505f0879e634e2a6c0b893cfe6bc133608be61197b7

Observation 50b6787f-db3d-4ff6-a9c9-ceeef55efee3 · outbound

This paper cites Fish audio s2 technical report.

dots.tts Technical Report Fish audio s2 technical report

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.054829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:c2874741813d1123586a0aedd67b4baabeea33008a7dc3c3f69b68c6d2783d87

Observation d62cfc91-7398-45ee-acda-3dc9648f1de1 · outbound

This paper cites MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis.

dots.tts Technical Report MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.051412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:88af8992779eae5e46c071f8c6371cf511fe76f2a108d4d92172c627c689934a

Observation da31b60e-9d12-4c5a-ab33-74f6270c59ac · outbound

This paper cites XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs.

dots.tts Technical Report XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.081577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:fb3665850a97d76ab7d305684221a3ff1afbe6c14b090875560b9b22c3357b33

Observation 951413b5-1f55-4636-91a8-be5539f7f333 · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

dots.tts Technical Report WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.076261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:fe6e6cc93a984d10912aaeef3d4f93b02adbb68add7eac4a60b7e319d03eab32

Observation f1567a9c-9f66-4bd0-be6c-4c6bf270299a · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

dots.tts Technical Report Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.070722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:b1345a582b6ceae350ca2d9c9dd3350285001ab1931eba017ef6e3a6b4c5f381

Observation a2908299-c8be-4a81-99a4-116b80418312 · outbound

This paper cites Ming-uniaudio: Speech llm for joint understanding, generation and editing with unified representation.

dots.tts Technical Report Ming-uniaudio: Speech llm for joint understanding, generation and editing with unified representation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.067246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:0461785bacfc85239e1e69120bc3b8fc44fa77b25a11fbc2c3b93bc620527209

Observation 293b4a13-3d26-4be9-aec4-7e6fc43bdeff · outbound

This paper cites vllm-omni: Fully disaggregated serving for any-to-any multimodal models.

dots.tts Technical Report vllm-omni: Fully disaggregated serving for any-to-any multimodal models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:57:20.065221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:44222fa8c73cba05f52b8cfdd0fd310ca1721073944efaa458039a387942b439

Pith citing papers

Observation 7e2e26bd-231f-4cdb-bcc0-08895295bebb · inbound

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens cites this paper.

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens dots.tts Technical Report

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T08:35:51.971819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:35:51.971819Z digest=sha256:d8c4b32f8b141d31df5191bba352f5175ae3dab6f35542e9acf17e430a7b911c