Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:56.850541Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 3 inbound Pith citation observations for arXiv:2505.19669.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:56.850541Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:54.536075Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T11:22:28.361623Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3db29749-9a2b-4027-bf1f-044ab8826a08 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling <eos> <bos> 𝑦! <eos> 𝑦#<bos> 𝑦
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ca17757-54e9-4855-a3da-83635fb401c5 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Specifically, it uses a Transducer to convert the text into a sequence of se- mantic token in real time
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation abd7dcf7-bf08-4280-a845-e12a5cd2ab21 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Delete ⟨Bos⟩ Mechanism
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b157e099-92a6-4c5f-86fa-8feddde23ab5 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7359fb2-2924-46d9-ac34-68d8b76debfa · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling 𝑚# 𝑚$ 𝑚% 𝑚& 𝑚'…… 𝑚! 𝑚% 𝑚
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 10fe7834-ee62-4958-a01f-13f37865905d · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Training Datasets We train SMLLE on the LibriSpeech dataset [27]
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 637d4ad4-31c2-4e52-a91e-db1f8143081f · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling SMLLE-R5
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 660d2a3d-5bc9-4dc2-b223-71d37dbf1aaa · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling It uses a Transducer model to convert text into semantic tokens in real time and reconstructs them into mel-spectrograms frame by frame using an AR model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e61516d7-dcd4-4687-81c3-8ab1b49cc20a · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling GPT-4 Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10855297-8798-4096-83be-7c47490509e8 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3a2c76e-35c3-4c52-be70-1536443c2ff4 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Zero-shot text-to-image generation,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91e8bcbc-9c99-4d09-9708-81e495670817 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Learning transferable visual models from natural language supervision,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8688be3c-dae8-4999-810d-adf59919d480 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fab4a743-7c6c-42bb-a1ab-0ef3ae3e293a · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5d276c2-d1f9-4b0d-b676-8d33a54b07dd · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Autoregressive Speech Synthesis without Vector Quantization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba626b19-461e-4a04-8d57-da15f014d083 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Moshi: a speech-text foundation model for real-time dialogue
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a56efe66-560a-4033-8703-cdd811a4327f · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 132089e7-bcf9-4ff1-ada2-857a733acffe · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bc85ce0-590c-4bb0-9971-49c256c354fc · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling SALMONN: Towards Generic Hearing Abilities for Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5c68240-e06e-4395-a491-797cb4e11261 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f9d2bf0-68e8-41e4-8fa5-9059d48c1758 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling WavLLM: Towards Robust and Adaptive Speech Large Language Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 943ecc01-d9c2-4724-a43d-1676c5147316 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0f38e56-6acc-4499-aa22-b0ff256ed87f · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Speech-t: Transducer for text to speech and beyond,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d76bc8fa-d7af-4976-923f-a96854527548 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Transduce and speak: Neural transducer for text-to-speech with semantic to- ken prediction,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d8ca4da6-6379-4673-9cd4-4329b7ea5d42 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5d1a7e1-0a42-4b77-9bae-a43bb0475ee4 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 042c7359-715b-468c-98a5-c2a1e3645329 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b835ee00-fb8d-4e72-8092-748c120d4af8 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ae8e0e9-d909-44bc-b465-2e2e7fd95af7 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb0d1b7d-887d-45c5-a7ff-2646d3a864b2 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9181c0a4-cfee-4a30-b9b8-fe8424ceb737 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling CLaM-TTS: Improving neural codec language model for zero-shot text-to-speech,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6446f42-6a4d-4da2-8133-f9a06b7a46a7 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01897a29-0ba8-4177-8598-5b7d749b6ec3 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling V oicebox: Text-guided multilingual universal speech gen- eration at scale,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 187165c8-e0e0-453a-9c32-094d0596e4c6 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Lib- rispeech: An ASR corpus based on public domain audio books,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4f863876-8cb1-4ca8-a1ed-7fe38cb5bdd7 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 041daa36-4dc0-4883-aa0a-bc187e0c0483 · outbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Zero-Shot Text-to-Speech from Continuous Text Streams
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b157e099-92a6-4c5f-86fa-8feddde23ab5 · inbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad5cbb77-9e8e-4f73-ad1a-344020d4bdf4 · inbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4706dce0-0b29-472a-80c5-38080d13a8c1 · inbound
Next Tokens Denoising for Speech Synthesis Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.