Pith. sign in

Paper Citation Record · LEDGER

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

As of 8 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 3 inbound Pith citation observations for arXiv:2505.19669.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19669 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:56.850541Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:54.536075Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T11:22:28.361623Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved22
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3db29749-9a2b-4027-bf1f-044ab8826a08 · outbound

This paper cites <eos> <bos> 𝑦! <eos> 𝑦#<bos> 𝑦.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling <eos> <bos> 𝑦! <eos> 𝑦#<bos> 𝑦

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.735947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:54.302540Z digest=sha256:93b504beff13186432dc1a2a7513fc16a95316b234f1af2d0b748ab7cf28f0ba

Observation 1ca17757-54e9-4855-a3da-83635fb401c5 · outbound

This paper cites Specifically, it uses a Transducer to convert the text into a sequence of se- mantic token in real time.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Specifically, it uses a Transducer to convert the text into a sequence of se- mantic token in real time

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.561077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:54.368789Z digest=sha256:49c64384a98c88708bde012cb15db9fe039642fb9f21dea27904eb56d0aa7d94

Observation abd7dcf7-bf08-4280-a845-e12a5cd2ab21 · outbound

This paper cites Delete ⟨Bos⟩ Mechanism.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Delete ⟨Bos⟩ Mechanism

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.399535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:54.457220Z digest=sha256:a61e92f9f24e86c8f87bd13d45f9d8ef4bbcbb144b1c2b7e1088a6cc45e876c6

Observation b157e099-92a6-4c5f-86fa-8feddde23ab5 · outbound

This paper cites Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.536075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.536075Z digest=sha256:2ddeef5a58f73b8cd0e8988e3d3f5d9c5141002dc8af2e17ba1d9ac48ce4a58f

Observation b7359fb2-2924-46d9-ac34-68d8b76debfa · outbound

This paper cites 𝑚# 𝑚$ 𝑚% 𝑚& 𝑚'…… 𝑚! 𝑚% 𝑚.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling 𝑚# 𝑚$ 𝑚% 𝑚& 𝑚'…… 𝑚! 𝑚% 𝑚

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.102532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:54.612074Z digest=sha256:47f7f132e3e13a5f48a9c668c1fd740c03c4d2a8a74813a1b8ce26d8642e8deb

Observation 10fe7834-ee62-4958-a01f-13f37865905d · outbound

This paper cites Training Datasets We train SMLLE on the LibriSpeech dataset [27].

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Training Datasets We train SMLLE on the LibriSpeech dataset [27]

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.849141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:54.679034Z digest=sha256:cd66dde0aad51716470eaad07d944c519622d7813ed79f8170f1bfd2cbfc9c56

Observation 637d4ad4-31c2-4e52-a91e-db1f8143081f · outbound

This paper cites SMLLE-R5.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling SMLLE-R5

Reference 7

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:15:59.537632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:54.734281Z digest=sha256:0eeb35026033e1060bb484b8e48f4e82e07a6892ecd497abaf5ce32a36641d88

Observation 660d2a3d-5bc9-4dc2-b223-71d37dbf1aaa · outbound

This paper cites It uses a Transducer model to convert text into semantic tokens in real time and reconstructs them into mel-spectrograms frame by frame using an AR model.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling It uses a Transducer model to convert text into semantic tokens in real time and reconstructs them into mel-spectrograms frame by frame using an AR model

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.284368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:54.800671Z digest=sha256:c207dece6a287c5e92f4ef6096e26c6ffd1463f875e3ed7d2dad5e74a8496f8d

Observation e61516d7-dcd4-4687-81c3-8ab1b49cc20a · outbound

This paper cites GPT-4 Technical Report.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling GPT-4 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.839595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.839595Z digest=sha256:91497e4591f860224d28b6d509fb05855baccd8ba07ea0d3292c59f0d73fe1ff

Observation 10855297-8798-4096-83be-7c47490509e8 · outbound

This paper cites The Llama 3 Herd of Models.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.910526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.910526Z digest=sha256:8be9d32f4fcd151682ad2384740342e42b739fd4133af2b80da49c3c10e76f55

Observation e3a2c76e-35c3-4c52-be70-1536443c2ff4 · outbound

This paper cites Zero-shot text-to-image generation,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Zero-shot text-to-image generation,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.998181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.998181Z digest=sha256:b3b9bf58f30516b304946924f278f7abf163ada2f1634bd6811de3cdeea3504e

Observation 91e8bcbc-9c99-4d09-9708-81e495670817 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Learning transferable visual models from natural language supervision,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.105958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.105958Z digest=sha256:f0b5a60b8c51b85508aabe6b6ca65e18e41242860a392e687a33cb62426f7cb3

Observation 8688be3c-dae8-4999-810d-adf59919d480 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.172894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.172894Z digest=sha256:f0fbcc41e67042f40c09053a3ba47c3f6bbc2dd6fefd176b3d3633c96e37bf3e

Observation fab4a743-7c6c-42bb-a1ab-0ef3ae3e293a · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.227783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.227783Z digest=sha256:4464e069f6e78bc44a43cc1b1f7e497d4f7fb0e7bdf62c94cfc5b591746c4ffa

Observation d5d276c2-d1f9-4b0d-b676-8d33a54b07dd · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Autoregressive Speech Synthesis without Vector Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.347789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.347789Z digest=sha256:d7ac44d0d2105043d37974ad9b89abe01e86ce26bd2cd3affe1b2bbfbc4e84b1

Observation ba626b19-461e-4a04-8d57-da15f014d083 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Moshi: a speech-text foundation model for real-time dialogue

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.424547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.424547Z digest=sha256:9d85cc40c6f48bc21956f7f70d6751011d6f7d9bd2eb2cd883e91654e7303057

Observation a56efe66-560a-4033-8703-cdd811a4327f · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.510988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.510988Z digest=sha256:102d620b78d110ccd70f5bc30cc16dbc026fb896ac8d42a7e8bdb40173cf1283

Observation 132089e7-bcf9-4ff1-ada2-857a733acffe · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.598942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.598942Z digest=sha256:de2ecd9d48ab6d5e03799aaed53390587d5ad7b59b68dfbaafc1b689b13b67a2

Observation 6bc85ce0-590c-4bb0-9971-49c256c354fc · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.630686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.630686Z digest=sha256:b806c44506a8bfb025ed95a6976fcad1e2026a710492bd994acccd773543e9a7

Observation b5c68240-e06e-4395-a491-797cb4e11261 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.684869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.684869Z digest=sha256:ce1998552edb55f53a1b4a7586e6b18ef2a22345d09c2390fc99bcfc7dda6618

Observation 0f9d2bf0-68e8-41e4-8fa5-9059d48c1758 · outbound

This paper cites WavLLM: Towards Robust and Adaptive Speech Large Language Model.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.743399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.743399Z digest=sha256:c41e65f55df79a433600e5447029f2d9b40da90c467ab30cd6fe736ba6a305a6

Observation 943ecc01-d9c2-4724-a43d-1676c5147316 · outbound

This paper cites LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.806875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.806875Z digest=sha256:88c73436850ca50ae80dee3ae8b883d461df67e5f5a5c735a776ccf9b0a39dd2

Observation f0f38e56-6acc-4499-aa22-b0ff256ed87f · outbound

This paper cites Speech-t: Transducer for text to speech and beyond,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Speech-t: Transducer for text to speech and beyond,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.070271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:55.885997Z digest=sha256:411d63603ea336d1e5c2928eff3265cfcc8ae2f40e89154f34e7eb2a2f2b698d

Observation d76bc8fa-d7af-4976-923f-a96854527548 · outbound

This paper cites Transduce and speak: Neural transducer for text-to-speech with semantic to- ken prediction,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Transduce and speak: Neural transducer for text-to-speech with semantic to- ken prediction,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.838464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:55.955500Z digest=sha256:732763b24a9f07de8be9cea380f10fd1303162522b3030a4003f09b67f86117f

Observation d8ca4da6-6379-4673-9cd4-4329b7ea5d42 · outbound

This paper cites High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:15:57.452738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:56.035770Z digest=sha256:ab855f0dd67f7a570f8fe249115a5d29f5145ee6c42c1c4099fdf58fb4ba43da

Observation a5d1a7e1-0a42-4b77-9bae-a43bb0475ee4 · outbound

This paper cites TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:15:57.238904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:56.168758Z digest=sha256:2fa8d96ef2f2221877913a73d3c51c30617894554cd85f2cacb0e6c5132b78f8

Observation 042c7359-715b-468c-98a5-c2a1e3645329 · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.248113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.248113Z digest=sha256:385386ecb72f3119e59269366a0178db71cf5bbdc3f5ace6509853a0574436b2

Observation b835ee00-fb8d-4e72-8092-748c120d4af8 · outbound

This paper cites ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.310389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.310389Z digest=sha256:c83383ffbfe260beb025d562471c3b9ed5d28590ab483bf12036ab8a67f69fd8

Observation 2ae8e0e9-d909-44bc-b465-2e2e7fd95af7 · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.366851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.366851Z digest=sha256:4db67af66f01be169849b8b8b3f11e90cb8c48bff36eab8ff12f0c40408299d1

Observation fb0d1b7d-887d-45c5-a7ff-2646d3a864b2 · outbound

This paper cites RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.438913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.438913Z digest=sha256:320bf3615109205a07442e58e478141278400c85cc7b6820be4bfa1fd22b0d5f

Observation 9181c0a4-cfee-4a30-b9b8-fe8424ceb737 · outbound

This paper cites CLaM-TTS: Improving neural codec language model for zero-shot text-to-speech,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling CLaM-TTS: Improving neural codec language model for zero-shot text-to-speech,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.542504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:56.518067Z digest=sha256:1023bde455bcf90d48f1eecbfc17fb1876a99662682d5847dfb89ae106fea006

Observation f6446f42-6a4d-4da2-8133-f9a06b7a46a7 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.602314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.602314Z digest=sha256:d6dbc2bee80db03f0e46ffaf12cd162a3de9e8f9522705a5b3ee4fd579580ac8

Observation 01897a29-0ba8-4177-8598-5b7d749b6ec3 · outbound

This paper cites V oicebox: Text-guided multilingual universal speech gen- eration at scale,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling V oicebox: Text-guided multilingual universal speech gen- eration at scale,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.207909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:56.695232Z digest=sha256:58504bced323411544e20014dcff18aebce2bf3001ea1e821ee6f9e850b2b1cd

Observation 187165c8-e0e0-453a-9c32-094d0596e4c6 · outbound

This paper cites Lib- rispeech: An ASR corpus based on public domain audio books,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Lib- rispeech: An ASR corpus based on public domain audio books,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.035412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:15:56.756041Z digest=sha256:b06df4f7e1a56e79ce999599e60052444a1a2ce9d0ce639915871f2ebf488690

Observation 4f863876-8cb1-4ca8-a1ed-7fe38cb5bdd7 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.797055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.797055Z digest=sha256:8ed4f5c0385c16b924c0617a79ee5330dafedad34d38b4403cebd18cc73e1ef0

Observation 041daa36-4dc0-4883-aa0a-bc187e0c0483 · outbound

This paper cites Zero-Shot Text-to-Speech from Continuous Text Streams.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Zero-Shot Text-to-Speech from Continuous Text Streams

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.850541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.850541Z digest=sha256:89919f229663b4d6854e05ff02fc045fc33ddce8a50d768447f940415c99ee4a

Pith citing papers

Observation b157e099-92a6-4c5f-86fa-8feddde23ab5 · inbound

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling cites this paper.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.536075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.536075Z digest=sha256:2ddeef5a58f73b8cd0e8988e3d3f5d9c5141002dc8af2e17ba1d9ac48ce4a58f

Observation ad5cbb77-9e8e-4f73-ad1a-344020d4bdf4 · inbound

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling cites this paper.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:02.492729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:02.492729Z digest=sha256:a50e7dd0f29ca7d67f8d4fcf0ba1960042a4eb98d433fdb8e4e3b9cf8d193469

Observation 4706dce0-0b29-472a-80c5-38080d13a8c1 · inbound

Next Tokens Denoising for Speech Synthesis cites this paper.

Next Tokens Denoising for Speech Synthesis Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:22:28.411634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T11:22:27.182862Z digest=sha256:228425fbae863d2155b9d1fa35f973822bf8ccf5b1cdb34ceca3be0203843dcb