Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T15:51:21.519785Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2605.27258.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T15:51:21.519785Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fc7704f2-7f34-4285-b233-67ebe54a3184 · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Neural codec language models are zero-shot text to speech synthesizers.IEEE T ransactions on Audio, Speech and Language Processing, 33:705–718, 2025
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebb4f618-57f3-4f8f-a8c1-bd1cd9384d8b · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Naturalspeech 2: Latentdiffusionmodelsarenaturalandzero-shotspeechandsingingsynthesizers
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a795d270-d0da-4fde-a878-3e70b2b63207 · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c95ab59d-13c7-41a6-8a74-498a687691d4 · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 58cd6d9a-78c7-44e0-9260-3917ea8b7bbc · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 347fa992-b235-4df3-881b-b67243fca0b3 · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4cc9a52-6e93-4129-aea2-5ad257ce243e · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e8a8ee4-dbb3-4761-9014-a1fa9575a1e6 · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Ditar: Diffusion transformer autoregressive modeling for speech generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a69d9ec-fdcd-48f4-96aa-a05ef31594e7 · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aff4cab-1178-49e2-983d-593eeb47c81a · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Qwen3-TTS Technical Report
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 15cdf836-4114-4b60-96f9-fb7dbc455bfa · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bae82a65-2eda-43a9-9008-40446c8fd93a · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Qwen3 Technical Report
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f3f5e993-4f09-45f9-874f-4060a0592f13 · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7ecd057-b6b1-4e10-9f0c-da38a15929c4 · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Flow matching for generative modeling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a88b261d-d9b4-4570-9773-fd17f1e1a673 · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Scalable diffusion models with transformers
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f35dc089-a771-43f4-b9e9-0382b5eee44d · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Hifi-gan: generative adversarial networks for efficient and high fidelity speech synthesis
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67d633e6-77f3-47d0-b950-0e82a738897f · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Powerset multi-class cross entropy loss for neural speaker diarization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcbb7191-fd24-491c-89d4-f14d98b195ef · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis pyannote
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56b4bf75-aaf0-4e06-b077-49bab1a65fa4 · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Dnsmos: A non-intrusive perceptual objective speechqualitymetrictoevaluatenoisesuppressors
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 452882a0-345e-46aa-a953-ac45dbc28eeb · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Paraformer: Fast and accurate parallel transformer for non-autoregressive end-to-end speech recognition
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49876449-2efa-40d3-9570-e9d683424d94 · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis FireRedASR: Open-Source Industrial-Grade Mandarin Speech Recognition Models from Encoder-Decoder to LLM Integration
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8ca12b10-cfd6-4ff2-a840-3f98804c826a · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Robust speech recognition via large-scale weak supervision
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6faba947-c387-4596-9ea8-02984d657d98 · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Qwen3-ASR Technical Report
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 039fd3d2-cbc9-4dc9-98d2-9dbc676c9cf3 · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis 3d-speaker-toolkit: An open-source toolkit for multimodal speaker verification and diarization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 684c4825-ec9c-41b0-9aa5-b0c38bcec73c · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Cam++: A fast and efficient network for speaker verification using context-aware masking
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d84b5f9-2bb2-44d5-b47e-e8381a32b754 · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 18c22e38-a338-4a3f-ad34-9c54977133aa · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Towards efficient visual-language alignment of the q-former for visual reasoning tasks
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4114d78-16df-4592-b387-291799c7c0fe · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dff122c5-eb64-4103-b03e-871c18861c87 · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis V oxcpm: Tokenizer-free tts for context-aware speech generation and true-to-life voice cloning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6e69b4a6-473a-4562-947b-a0893f0fda82 · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis VibeVoice Technical Report
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f04624a4-aa91-43ab-8770-f7a9e344cf83 · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Fish audio s2 technical report
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ac48d7d0-7235-4016-ac85-5d345ec6882a · outbound
PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
No inbound Pith citation observations are available.