Pith. sign in

Paper Citation Record · LEDGER

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis

As of 6 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2605.27258.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.27258 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T15:51:21.519785Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact14
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc7704f2-7f34-4285-b233-67ebe54a3184 · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers.IEEE T ransactions on Audio, Speech and Language Processing, 33:705–718, 2025.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Neural codec language models are zero-shot text to speech synthesizers.IEEE T ransactions on Audio, Speech and Language Processing, 33:705–718, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-29T15:51:21.519785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:2f4c88c596eacd817b5908458e992b93cf29fcb92a5ade2e7ec299950e84b39a

Observation ebb4f618-57f3-4f8f-a8c1-bd1cd9384d8b · outbound

This paper cites Naturalspeech 2: Latentdiffusionmodelsarenaturalandzero-shotspeechandsingingsynthesizers.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Naturalspeech 2: Latentdiffusionmodelsarenaturalandzero-shotspeechandsingingsynthesizers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-29T15:51:21.519785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:8d27a4b746269ef11701dd2bd78e63256e708c124566158beae98cb8bb980e8f

Observation a795d270-d0da-4fde-a878-3e70b2b63207 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-29T16:23:39.877202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:82f34c4daf4188f6327c7a620e73988f183ccf71310963eb2e638c2d61228c25

Observation c95ab59d-13c7-41a6-8a74-498a687691d4 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T16:23:39.879632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:5f43a5108ff0be14f64b1fd1963efda9207a4e45c8fdc74225baf73a65b1db31

Observation 58cd6d9a-78c7-44e0-9260-3917ea8b7bbc · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-06-29T16:23:39.894622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:440baa54e2a2ec90fbfbc0f82d96b60c0ff4760db89985bd57a1f39475acbb3a

Observation 347fa992-b235-4df3-881b-b67243fca0b3 · outbound

This paper cites FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:23:39.902062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:978aaa8b6a36487ef08f1a12af9a6e89e9b64600ecc5f0240cbf4324d1eee7c6

Observation d4cc9a52-6e93-4129-aea2-5ad257ce243e · outbound

This paper cites F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T15:51:21.519785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:767a6e15143812698da7f81f974d8fc13ca0348176691a675f96d71db28bed62

Observation 4e8a8ee4-dbb3-4761-9014-a1fa9575a1e6 · outbound

This paper cites Ditar: Diffusion transformer autoregressive modeling for speech generation.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Ditar: Diffusion transformer autoregressive modeling for speech generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T15:51:21.519785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:a7ac352d5603507f4d30bbc967958feef019bebfb4544b58ac9c40feb61ce5ab

Observation 9a69d9ec-fdcd-48f4-96aa-a05ef31594e7 · outbound

This paper cites Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T15:51:21.519785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:f1c0a3a310ea3f62721c43160ff08e079ef48c5a4c264d5374f926eb7d6296ea

Observation 6aff4cab-1178-49e2-983d-593eeb47c81a · outbound

This paper cites Qwen3-TTS Technical Report.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Qwen3-TTS Technical Report

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T16:23:39.898947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:ca45197ec9de3279977f1040e2dcca6fb508c8f959739b1b857bea2825146b22

Observation 15cdf836-4114-4b60-96f9-fb7dbc455bfa · outbound

This paper cites MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:23:39.874719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:c18c2faf390553a88f3630910701d9cc36a8c7214d745a9ebfc036010d1ddf1b

Observation bae82a65-2eda-43a9-9008-40446c8fd93a · outbound

This paper cites Qwen3 Technical Report.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Qwen3 Technical Report

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-29T16:23:39.896818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:7ec8c6d4101c6cefc30c068a03e02ae4291d3a5df113a982549372bd0ac89ba7

Observation f3f5e993-4f09-45f9-874f-4060a0592f13 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-29T15:51:21.519785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:b096421ac9956121c18b1f03ea2333d5f57982658d1b54d7e1722ecb23488a01

Observation d7ecd057-b6b1-4e10-9f0c-da38a15929c4 · outbound

This paper cites Flow matching for generative modeling.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Flow matching for generative modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-29T15:51:21.519785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:644c83521b47d01b0c84648df946a20dba1fe7c4136486b079d772f354b5a7bc

Observation a88b261d-d9b4-4570-9773-fd17f1e1a673 · outbound

This paper cites Scalable diffusion models with transformers.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Scalable diffusion models with transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-29T15:51:21.519785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:b7c7914b875d049de92866f410285d31d24500f79f1c13074c8b0be1d764d697

Observation f35dc089-a771-43f4-b9e9-0382b5eee44d · outbound

This paper cites Hifi-gan: generative adversarial networks for efficient and high fidelity speech synthesis.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Hifi-gan: generative adversarial networks for efficient and high fidelity speech synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-29T15:51:21.519785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:011cfccb69aef223fe3238f4fe3b74d49483b9952922de648556df495632e492

Observation 67d633e6-77f3-47d0-b950-0e82a738897f · outbound

This paper cites Powerset multi-class cross entropy loss for neural speaker diarization.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Powerset multi-class cross entropy loss for neural speaker diarization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-29T15:51:21.519785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:689020a03b83529219757d1d109c3aa0c940081dcca1f0e39d5d432d877134e7

Observation bcbb7191-fd24-491c-89d4-f14d98b195ef · outbound

This paper cites pyannote.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis pyannote

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-29T15:51:21.519785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:ba1f37f37c686e86c0a85e562f06899fe66d9afdee0cf646e7c62afbcb8d3ef0

Observation 56b4bf75-aaf0-4e06-b077-49bab1a65fa4 · outbound

This paper cites Dnsmos: A non-intrusive perceptual objective speechqualitymetrictoevaluatenoisesuppressors.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Dnsmos: A non-intrusive perceptual objective speechqualitymetrictoevaluatenoisesuppressors

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-29T15:51:21.519785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:9244aedca170f04f820f97b61ef049d09c3dd5b7744e0a44dab0aa86257f0387

Observation 452882a0-345e-46aa-a953-ac45dbc28eeb · outbound

This paper cites Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-29T15:51:21.519785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:ef56842bf98e01a2507557cd0a570ae7ccbd29b594f84c2a573980efa5e30e02

Observation 49876449-2efa-40d3-9570-e9d683424d94 · outbound

This paper cites FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:23:39.871989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:d1db44d009d902be5eddcfad3d57eb6700e6982708416e2904a0f355871516b0

Observation 8ca12b10-cfd6-4ff2-a840-3f98804c826a · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Robust speech recognition via large-scale weak supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-29T15:51:21.519785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:1f0991ae7dc71b3cc8b0655e491498fde03f5d44112db7f906b13aae7d3a594d

Observation 6faba947-c387-4596-9ea8-02984d657d98 · outbound

This paper cites Qwen3-ASR Technical Report.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Qwen3-ASR Technical Report

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-29T16:23:39.892254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:a6c5236b3248c2e00f70e51d18e07c0ec3c7f0b12755b30b58b75b95e548cc21

Observation 039fd3d2-cbc9-4dc9-98d2-9dbc676c9cf3 · outbound

This paper cites 3d-speaker-toolkit: An open-source toolkit for multimodal speaker verification and diarization.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis 3d-speaker-toolkit: An open-source toolkit for multimodal speaker verification and diarization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-29T15:51:21.519785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:f92fe62a0aeabfed81e21652907a30a6d76b38c313f559320c14be98bb9b8e6e

Observation 684c4825-ec9c-41b0-9aa5-b0c38bcec73c · outbound

This paper cites Cam++: A fast and efficient network for speaker verification using context-aware masking.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Cam++: A fast and efficient network for speaker verification using context-aware masking

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-29T15:51:21.519785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:6743c31162a9528fb5af0ca6f80c0f40b5bcdd3a53c186932d96dfa55ba0d28b

Observation 3d84b5f9-2bb2-44d5-b47e-e8381a32b754 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-29T16:23:39.881950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:23cbe6665d7b2f05e16e8148d44e904d1b4397e8fd7f409322b76833b32d93dd

Observation 18c22e38-a338-4a3f-ad34-9c54977133aa · outbound

This paper cites Towards efficient visual-language alignment of the q-former for visual reasoning tasks.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Towards efficient visual-language alignment of the q-former for visual reasoning tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-29T15:51:21.519785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:378ee49526738129a93fbedf90fdc75296706bdf120a80c9a68d79586a3bd031

Observation c4114d78-16df-4592-b387-291799c7c0fe · outbound

This paper cites W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-29T15:51:21.519785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:ffcdfc3d6b25a7f9282d39cb5bd6912172b5cbd97400c95a0508575128fa0798

Observation dff122c5-eb64-4103-b03e-871c18861c87 · outbound

This paper cites V oxcpm: Tokenizer-free tts for context-aware speech generation and true-to-life voice cloning.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis V oxcpm: Tokenizer-free tts for context-aware speech generation and true-to-life voice cloning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:23:39.884569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:ec047c40aa5d10826e180f2ea91a4ee8a620c3dfc9bed7a70fb1e3f99b8407d9

Observation 6e69b4a6-473a-4562-947b-a0893f0fda82 · outbound

This paper cites VibeVoice Technical Report.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis VibeVoice Technical Report

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:23:39.904655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:7b30c3d9616303b2c9377349fe3aed777c0cb246cd57dbd6f56ca9a8431bbbe0

Observation f04624a4-aa91-43ab-8770-f7a9e344cf83 · outbound

This paper cites Fish audio s2 technical report.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Fish audio s2 technical report

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:23:39.887587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:7853c6be30bec7dde65831ad888041dcc0234eeda50aa98cac74c665f78977b1

Observation ac48d7d0-7235-4016-ac85-5d345ec6882a · outbound

This paper cites IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:23:39.890163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:6b186babec088983f27a5d753f32d47498f3bf940c763166dc96f1f525b57e18

Pith citing papers

No inbound Pith citation observations are available.