Pith. sign in

Paper Citation Record · LEDGER

SpeakStream: Streaming Text-to-Speech with Interleaved Data

As of 17 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2505.19206.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19206 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:22:31.869618Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:52:44.054140Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:07:12.867079Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c797376-1075-4f7a-b805-ed8753374b7c · outbound

This paper cites Audi- olm: a language modeling approach to audio generation,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Audi- olm: a language modeling approach to audio generation,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.551223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:22:30.401774Z digest=sha256:d8755b19115aa5560a4c9715ac18d41df845360a915594570537777167f656d5

Observation 2a49efeb-c160-4df2-b2ed-2eb489baef6e · outbound

This paper cites Qwen2.5-Omni Technical Report.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Qwen2.5-Omni Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.544240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.544240Z digest=sha256:0d8c680cb84a0063eae7795d0a511e9c0ffd35c68c85055ec39c5c00995fcf06

Observation bc980c59-7ca6-4cfc-9ad1-f31144368a47 · outbound

This paper cites Spirit-lm: Interleaved spoken and written language model,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Spirit-lm: Interleaved spoken and written language model,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.595293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.595293Z digest=sha256:0aa9552e1346402e54c37dc1050285952ace22e2214ee3609753f86c61381e5e

Observation c1bc2262-c193-473e-821f-5069f1f8eb6d · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

SpeakStream: Streaming Text-to-Speech with Interleaved Data MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.673231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.673231Z digest=sha256:09d5cc96f0c408be39595c16d7a4d3160c6f15dbd398785c22b6e595f873428a

Observation 0b239529-5d94-462f-9b54-9ad986574e10 · outbound

This paper cites Zero-Shot Text-to-Speech from Continuous Text Streams.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Zero-Shot Text-to-Speech from Continuous Text Streams

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.726922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.726922Z digest=sha256:6666b961a1558eec3858b478f85934d3be1dacb63756ed6661b010bd68ab1db0

Observation 8928467d-c253-4fba-ac9c-676f4539d528 · outbound

This paper cites Speak while you think: Streaming speech synthesis during text generation,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Speak while you think: Streaming speech synthesis during text generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.477254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:22:30.780606Z digest=sha256:7257bde141e973a1af17787596699d70f36019e360f4e01b9b6927e43ee27b68

Observation b18d51a4-4449-433f-b0d8-7f587de7c478 · outbound

This paper cites Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.842860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.842860Z digest=sha256:89141fb4b526364bda23e9860e9d5209d6bd43e4ca6450f775a5214725f24485

Observation c136ad76-0c3f-424d-97f2-a7621681b0e8 · outbound

This paper cites dMel: Speech Tokenization made Simple.

SpeakStream: Streaming Text-to-Speech with Interleaved Data dMel: Speech Tokenization made Simple

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.915758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.915758Z digest=sha256:5c19b7d8209a58bd3edd1049ed2c682b9b3cdff0993598ab75d7fc4e60f4ca08

Observation 4d648051-cf1a-46e1-9b8e-67188104c6c8 · outbound

This paper cites A 3T: Alignment-aware acoustic and text pretraining for speech synthesis and editing,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data A 3T: Alignment-aware acoustic and text pretraining for speech synthesis and editing,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.434512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:22:31.004924Z digest=sha256:598dda8d378775ee595b69fbfb1e50288b2d0a3306af4f4bd04bc770ac3aa355

Observation 9ffb4b08-22c6-47bd-ad27-01f64133f11b · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.029594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.029594Z digest=sha256:bb73a650c99821c600c3b0ecfb9345c2574ed20e4aedabea8fea0b6b4126685a

Observation 204d923a-49c9-443c-ad21-c47bea9e7cbb · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

SpeakStream: Streaming Text-to-Speech with Interleaved Data CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.087888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.087888Z digest=sha256:a15a3386cb833028f9c514ffc9c561a4a91a1d9cfb92e0149374eddb76edc4ee

Observation 69c1a5bb-5c73-4bd8-8a8f-4cba12c0e441 · outbound

This paper cites E3 tts: Easy end-to- end diffusion-based text to speech,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data E3 tts: Easy end-to- end diffusion-based text to speech,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.343081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:22:31.132497Z digest=sha256:2d5bbae0bba8577c561fb77c663b48729709c075ee768bf5922360c9985a2ce7

Observation 7fc88f8e-e561-4c13-89a6-35e5a654f82f · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

SpeakStream: Streaming Text-to-Speech with Interleaved Data FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.169600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.169600Z digest=sha256:a9799b846f289799b012810980dd36fbb44b67622e65f7f7b2f451d7974339c1

Observation 6fa8f085-6f5f-4e23-8147-693b3da14f4e · outbound

This paper cites Tacotron: Towards End-to-End Speech Synthesis.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Tacotron: Towards End-to-End Speech Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.221168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.221168Z digest=sha256:90d7db1dba8ada27eeab63ff3a914340b5c2a9b2a52dd0bc2e1c57dd25602cb5

Observation 30b3b16d-7403-4add-a2d3-9a216fec6c90 · outbound

This paper cites (2024) Text-to-speech guide.

SpeakStream: Streaming Text-to-Speech with Interleaved Data (2024) Text-to-speech guide

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.304761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:22:31.249785Z digest=sha256:20841d65d79b5225af5c7461aa096eae701a75773391f0b5a3472d3fd7306879

Observation 1efec5e5-675f-4cf5-8bed-cd3d766ca544 · outbound

This paper cites Streamspeech: Low-latency neural architecture for high-quality on-device speech synthesis,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Streamspeech: Low-latency neural architecture for high-quality on-device speech synthesis,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.252460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:22:31.282325Z digest=sha256:45e16235acf922510049809a6813cba97cd426b453a3ed1552c506df8f4359f6

Observation 6209cf8c-2fd9-47ab-b891-3aee83ad4a78 · outbound

This paper cites MLX: Efficient and flexible machine learning on apple silicon,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data MLX: Efficient and flexible machine learning on apple silicon,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.327328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.327328Z digest=sha256:7eb5e5f72dfa063138f213c5907bc31509749300ce249017f7681cd2d9b27081

Observation e261da5d-e116-4fe1-b3ab-5a9103f733a5 · outbound

This paper cites VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech.

SpeakStream: Streaming Text-to-Speech with Interleaved Data VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.416374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.416374Z digest=sha256:7bb6e627ba7175942ae25364a63d7bc5eb727677b4cc36c41084f03a5ef8d0bc

Observation 20ab1739-357b-457d-91d3-b16a217f4b12 · outbound

This paper cites BERT: A Review of Applications in Natural Language Processing and Understanding.

SpeakStream: Streaming Text-to-Speech with Interleaved Data BERT: A Review of Applications in Natural Language Processing and Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.501147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.501147Z digest=sha256:ae59a440f29037346cdb877fdc5da993f194b194139643646ad4d91a54ef2dd8

Observation d9bcda6d-b7a2-4e38-90e2-dbf56765bdbf · outbound

This paper cites Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.555872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.555872Z digest=sha256:dfe7b4d4f72b38012853018fd69f8312eae53da61d51edbf3ffa4269da7f3233

Observation 264a593b-6af6-4e94-a1a6-e63767c50fec · outbound

This paper cites LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes.

SpeakStream: Streaming Text-to-Speech with Interleaved Data LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.639465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.639465Z digest=sha256:70bfc9c5a8365b45f001242da062a4f358c149ad41cd781f7006f12a3141143a

Observation b70e1eb8-b561-42d4-bcc2-511129d73975 · outbound

This paper cites Transduce and speak: Neural transducer for text-to-speech with semantic token predic- tion,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Transduce and speak: Neural transducer for text-to-speech with semantic token predic- tion,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.190432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:22:31.674377Z digest=sha256:ea71764c9f5422219945df0d6a29c119f9bb9da49bc070e809d7603c89c68256

Observation fcf3c1f7-9d52-4c8b-b2ef-cf855d75e533 · outbound

This paper cites Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.715642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.715642Z digest=sha256:3130800851dc5224ef145b0a77afdacb54444c137fe1c05e3bf6456f6ab8b9f9

Observation 5fc08e52-2a4b-4dfc-88ac-3e787debc790 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.733397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.733397Z digest=sha256:2c522cb760eb3557923eb41cc52b91f77f01fdeb991488398ebf4a02b14c7394

Observation bf912353-3894-4d2b-86e5-1a360d45569b · outbound

This paper cites Bigvgan: A universal neural vocoder with large-scale training,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Bigvgan: A universal neural vocoder with large-scale training,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.075120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:22:31.746547Z digest=sha256:593255300124c5549c9505b57b37337f61e5580d769c741a6a22d8371ad1e5da

Observation 939b6549-48d9-4fc6-b00e-96a60271dc9c · outbound

This paper cites V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.020489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:22:31.759386Z digest=sha256:fd60d3d66c28266078e0df8a15480213d39fd3de993fe695a798e2a09f0cd362

Observation 027f9df9-6925-4d0d-941b-fc979cc4d52a · outbound

This paper cites Non-causal to causal ssl-supported transfer learning: Towards a high-performance low-latency speech vocoder,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Non-causal to causal ssl-supported transfer learning: Towards a high-performance low-latency speech vocoder,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:32.945758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:22:31.778855Z digest=sha256:de44e8704c14043ad44e23b01d96db73e1d708fc5c31368c85ddf20b818d6dd5

Observation 64393298-7ab5-4cbc-933d-fc9615b7d845 · outbound

This paper cites Coqui TTS: A deep learning toolkit for Text-to-Speech, battle-tested in research and production,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Coqui TTS: A deep learning toolkit for Text-to-Speech, battle-tested in research and production,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:32.906779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:22:31.787676Z digest=sha256:35796c7f294196f78735ce5d9b09e00c30995c2183f69833d93342dd6aa35993

Observation 71b7ac5a-074f-4634-8442-ef81b3ea1d29 · outbound

This paper cites The lj speech dataset,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data The lj speech dataset,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.807757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.807757Z digest=sha256:038e0921e0314d9fae8f87d4b0134575e793c5cecfaf0979a932408847724a40

Observation 3fd2acba-0c0e-4856-a3d9-9b89e5eb8477 · outbound

This paper cites Whisperx: Time-accurate speech transcription of long-form audio,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Whisperx: Time-accurate speech transcription of long-form audio,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:32.766258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:22:31.816394Z digest=sha256:1c27cec57f62301d81e96a871337fba8f4f8190a60c416c72efb4ce20301f2fb

Observation ad6a3e0d-0db1-425d-96dd-3d850d4678a6 · outbound

This paper cites Robust speech recognition via large-scale weak super- vision,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Robust speech recognition via large-scale weak super- vision,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.823985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.823985Z digest=sha256:5675aba270a1dc3e807e693e1ee90174ddc3cd055f0795f5e55dc14f7adce82b

Observation 44a264d3-1988-48c0-bc65-f38876f91a97 · outbound

This paper cites Libritts-r: A restored multi-speaker text-to-speech corpus,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Libritts-r: A restored multi-speaker text-to-speech corpus,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:32.674771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:22:31.845205Z digest=sha256:0af91e221c5ab17083d04a2d30c8dae59894fcfb1afe36d06cf4c52786858285

Observation f1fbf252-58e1-4706-91c0-32a9c352fb57 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Librispeech: an asr corpus based on public domain audio books,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.853069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.853069Z digest=sha256:aade677b3d68cb1cd0320f0fb3f8e7dbcfcda87bfce848198f327a0d73dbf02d

Observation 841334ad-98d8-4578-a9aa-8572b38c2b0b · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Moshi: a speech-text foundation model for real-time dialogue

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.869618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.869618Z digest=sha256:7ffe3fb41eb928035f87508ab5ece8ebb2048c35734023a14491b2bc671a439e

Observation d2eae670-95c0-493b-8556-5d1d38b0af9b · outbound

This paper cites Available: https://www.coqui.ai.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Available: https://www.coqui.ai

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:32.851116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T14:22:31.800042Z digest=sha256:bc604567de44cc1f4b88f366a5a025a7d136836252183a1a2055586ebda310d1

Pith citing papers

Observation 5fea5dc5-c540-4237-bc1d-d1f003ccf490 · inbound

ChipChat: Low-Latency Cascaded Conversational Agent in MLX cites this paper.

ChipChat: Low-Latency Cascaded Conversational Agent in MLX SpeakStream: Streaming Text-to-Speech with Interleaved Data

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:44.054140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:44.054140Z digest=sha256:62459ce32ef7a030f91d7c84d04caaf6997abec7c7d58a83375bb3765ed5d519

Observation 6227ce77-3ec2-40de-9470-32c1250b84ee · inbound

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding cites this paper.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding SpeakStream: Streaming Text-to-Speech with Interleaved Data

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.868675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:cb8b267805490714f247597952c26b63f8bd3d5be4d9e158f51a34cf18464900