Pith. sign in

Paper Citation Record · LEDGER

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis

As of 22 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 3 inbound Pith citation observations for arXiv:2412.16102.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.16102 v3

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:50:49.092653Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:56:54.401206Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:52:04.433992Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5abad7c9-a9d2-4dfe-a222-0e9041616592 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:48.786023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:48.786023Z digest=sha256:2c9d74c473ea15d186524705d0e07e95a361588f426858f24e82a8e1a6c035f1

Observation b1fe770c-7af2-4e8f-85e9-86fa99ccdc68 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:48.792135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:48.792135Z digest=sha256:1680303299dad2fb8b37e2c5b9477e039b766ae7154687ffd8f4adaab6b7c736

Observation 608e2db3-527c-4222-8601-e0d3c92feb6b · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:50.246432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.798263Z digest=sha256:cbfe04868c7144a18b4ff16d8616ff6d30599ee7f0cbc2ae496642da9ac67330

Observation 68bac971-2d16-4aa3-b4b2-16674cbe49a3 · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:50.230161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.803566Z digest=sha256:a44d4b4949469a9783e5092f46ff7ab2ebeeb411162a7b270f9ae8f1606c7443

Observation a85879f1-7564-47f2-8945-97aa42eeac38 · outbound

This paper cites Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:48.808748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:48.808748Z digest=sha256:69c5c8efa5424850997e2cf354d9e04fd5f5ad7a4c04e1a5fb7f3c091a192f4b

Observation f4148da6-3c4e-4d9f-bdbd-5b7ac08b07af · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:50.213563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.814243Z digest=sha256:a0aec0fac102e516fd51a8163323af71a184255a3a2eb9499dd08198b378043e

Observation c953c41b-7e6d-4422-8d8f-dbe3ef932164 · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:50.195284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.820206Z digest=sha256:280e94924e61f97a64adeda0729988b0c60ee0800acc1cb361ff8b4e62cc1ff0

Observation d7e6913d-f226-4a9e-9dc8-f132b54e1678 · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:50.179158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.824949Z digest=sha256:a308976a3f81e96f61e2a5361a239eb6e058815aec08e4486bc32553b3ec4bb8

Observation a27e6bf5-dd16-48a4-b502-0299cba8771c · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:50.161664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.830203Z digest=sha256:615a9c74afc0aabb10513be4b5fc4536e7cc25ae4a9c5726dc9f2d25b5ad025a

Observation 3f5245a7-f4a1-49f4-8297-a9b7af24af27 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:48.835195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:48.835195Z digest=sha256:8995ce604686d4d191730bdc1e7170685b33c90762426abd08ba96c338946fbd

Observation 99cc4685-422a-4cd0-91c2-c008420c40ce · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:50.144730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.840660Z digest=sha256:25a291e719e2e5a80b2725190a20003bc2441d656076b3738b802c6f54a8d913

Observation 2b146d94-fd67-4d8a-a121-1cba0b0853a8 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:48.845673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:48.845673Z digest=sha256:d7fefa9554813e893b75b2503268ca4ec3c7f6e53dadd5cc7c36fd859b7856f9

Observation 49863daa-bcf8-47f1-b54b-0375eb04efe1 · outbound

This paper cites Assael, Brendan Shillingford, and 1 others.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Assael, Brendan Shillingford, and 1 others

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:50:50.128631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.850883Z digest=sha256:8174385ebb82dedc5a63453e66ed93e390c97775ff3e43a978ec1f844e924dae

Observation a3be91f4-229e-4d20-bd13-2aea0368aa51 · outbound

This paper cites LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:48.856109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:48.856109Z digest=sha256:e37fac838f22a32abd6671e8dc3aec58270c18a8cf3792779d2e2bdd00d05652

Observation fa140664-ae2b-4f37-9c28-9570a3db10ad · outbound

This paper cites Zero-Shot Text-to-Speech from Continuous Text Streams.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Zero-Shot Text-to-Speech from Continuous Text Streams

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:48.862070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:48.862070Z digest=sha256:43e8af83ad2ecebd3a315877d84552953ca8fb31c4001381003aac48d160f085

Observation 904b4ade-b7cf-4e5e-a578-f8fe0250750e · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:50.111401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.867310Z digest=sha256:3e655aa4d6217c9768214b1b69f8de1710e392fd493831e75cd5ae6416a150a6

Observation 90863e13-c9b8-4101-861d-0d5762bdc042 · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:50.094937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.871942Z digest=sha256:639abf8f3a277b29189859d48773875ec103617ba83fa6b04b27ca5efa69563e

Observation e9e6cbd4-2ac1-4c1a-99b9-6f5ca1aa60a4 · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:50.076386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.877046Z digest=sha256:19b4f4bb3f6b7e36bd57482c3e496a91113751e1d632472dc62dd3cc2cd8ee17

Observation 5e1f0168-0105-408f-bc37-96ad5c19245a · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:48.882325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:48.882325Z digest=sha256:6abbd60bb6bec4fd9923808b94255de858a4b59594fa125a792304c0ae071d34

Observation 73d3a66c-a74d-4113-a959-0ae78c8ec21d · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:48.887742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:48.887742Z digest=sha256:e1b8d6c7cf28097487d7a634683401a60b3234f988e51de8da58e4116c9b8cec

Observation d8163073-2598-4678-8751-2845f62ee83f · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.972137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.893463Z digest=sha256:3c598d2a8bf9e3d1c6a1aa42e63012877604332e473db7e4d69b9ab0d3d48d49

Observation 76646904-ff98-4c32-b4a0-1f8f9001f0af · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.955051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.899149Z digest=sha256:01af46ffc2fad40374837ac57e966c762c5b98e5f4ebb23b8d6973cad128c46b

Observation a09c2152-1153-4bc3-b51c-7e98aaefb144 · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.938269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.904239Z digest=sha256:f9a0898d673ab7d69a80d4dde7d4594ece809890cf0832bc4e204c5058aeb1f8

Observation 3cefdc99-1888-47a2-8bfa-07772a00ec2c · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.921675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.908958Z digest=sha256:725b2e1a3d7451a2ec4265bc063694ac3fba741ad09aa317f3a09ca2a8f0fd47

Observation f526b32d-ca66-4533-9cf7-164a3f0e8e8d · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.903323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.914434Z digest=sha256:ceb562a76e411ae1b878e76c315be8933cdd12931b902c834c0436a78770978f

Observation 4db010cb-b496-40ba-8481-68349392e9c0 · outbound

This paper cites Weiss, and 1 others.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Weiss, and 1 others

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:50:49.884580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.919212Z digest=sha256:637173e2c4a5c737d1f3be066b9f2acb3ab2b2a72c8557db9127d199583b6643

Observation 95777c88-558a-42a0-ad13-1da4a7653408 · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.867196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.925636Z digest=sha256:5d6fe1cb9028a19fceb7036730e429fc67c0df055a7ac821a3ef8e68d41411ec

Observation b9c2baac-c41f-4619-89d2-7fd1ab8300bf · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.850436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.930238Z digest=sha256:8d09f9007d4767a12dad9eaeb4e96c208edd5444902e5f03c6cbe97d28492aa5

Observation 68c1d559-5220-4df1-83cd-6310dbe9cfbe · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.833197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.934915Z digest=sha256:83e9dd9a2a619edcc02e36b9de68e21e47d0258f87e0bbfbc2e6f131b04681f8

Observation b8f29230-c0fb-4d4c-935f-968d993dab08 · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.815539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.939655Z digest=sha256:a59d7c0401c0420b5d5edef1b4553c71300a9a3abb976eef585df500d21b2a69

Observation 3c35198d-3d1b-4137-8afa-9f31ce5e0121 · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.798705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.944674Z digest=sha256:dc11fababb9f56f6ba51b6c76d9385e6807320bc7e5854166ad1b7c952a2a230

Observation 0700a401-29e4-4052-8d74-92c5d4a6da43 · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.781067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.949568Z digest=sha256:44536fa6e1f14da0f7486fa09cad1e467559f264796b520ed911fc194339bdbd

Observation 2020cdbb-4002-4a58-8b59-cb0e193aea04 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:48.954317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:48.954317Z digest=sha256:0a8d0d0a12ae0ca9ddb35a5bf7be111c2831e2f108aea0f6e22991a6d9b939e8

Observation 30cf56df-974d-45e8-a5af-eb694d606638 · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.764752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.959404Z digest=sha256:bcc735843132fa16aa13df12ea51a0ed83b8da939ab599527b3c1113ed378244

Observation 2f843f18-b863-40f1-abeb-60f527d6c75b · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.746943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.964574Z digest=sha256:155afa47b6086fbb751e982f574da50ede66b619b55c9ff39031b93bc1b28a63

Observation 8b799533-3325-4bde-8e14-8d9bdcf774ef · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.729821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.969402Z digest=sha256:0f6fb9d28d44759d531f8fe448b7ece658a616074d25b332e5693891905579e7

Observation 02341cbf-8943-44cd-bf2d-b7c3eb318834 · outbound

This paper cites An Embarrassingly Simple Approach for LLM with Strong ASR Capacity.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis An Embarrassingly Simple Approach for LLM with Strong ASR Capacity

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:48.974260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:48.974260Z digest=sha256:ddf02913eed73a710d3578f30a5af9ae7f541c0061de990283d97d74655a3756

Observation 41dacecb-7a32-48e5-bd4a-bd2f39fe52eb · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Autoregressive Speech Synthesis without Vector Quantization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:48.979014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:48.979014Z digest=sha256:dd29e660f612b3545af17583940192afdfbc2340d786829a0d2de64d1d98148f

Observation e9ab39a5-cec5-4a91-aef3-9d769a3a6eb1 · outbound

This paper cites Spirit LM: Interleaved Spoken and Written Language Model.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Spirit LM: Interleaved Spoken and Written Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:48.984452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:48.984452Z digest=sha256:4cacc82220e285d492f7b2f5e0ef0800c17e49bc36409118acb0b570841b0329

Observation cbd5eb9f-a4ac-4791-babe-43352f2a4842 · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.712046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.989712Z digest=sha256:8251b2ae59d535acf394c05800305bd9114fe2992ea0e64488e171c8cebfacb7

Observation e4210081-0b81-4d23-ab2d-7120bbee6044 · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.694273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.994506Z digest=sha256:0d0776faca67866acb40f35a1437cf56108d4c71455419fd0f7410fa8d12126e

Observation d01f23f2-e972-4f4b-b657-dbbc12cc4cce · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.676068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:48.999266Z digest=sha256:fa77db7e1e04c054ceb2de3d9ec0b91cbd954d8d23a0c74ac132201df7882d00

Observation 827437ce-fad0-4902-b85c-90aa0079c1a8 · outbound

This paper cites Weiss, and 1 others.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Weiss, and 1 others

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T10:50:49.657099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:49.004549Z digest=sha256:1a6ae4c1b5a13f09d7eca32fe2826d42d4dc0623ad588c02d5bebf564feae158

Observation 06bef987-0894-470b-9668-add4ed54f302 · outbound

This paper cites TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:49.009679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:49.009679Z digest=sha256:f7ceb74e6781ed6cb40700b10a83a8ec47a6b4a9cbaa1ccee8d98afdce71d7cf

Observation a9d98538-0be9-4d7b-bbdb-f1635063d93a · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.638887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:49.014562Z digest=sha256:9bc49197851d39a4c3f408464ffcfb9b2a72384b97dc68f121d894de2722b103

Observation 0447ea0a-cbd9-4840-9ce2-69080204b639 · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.621392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:49.020300Z digest=sha256:55cb78b89bf9ef39c028696db8ae312a4b52c8416ed87edbb5008ba62f43a5de

Observation 1a01b47c-a422-473e-ad87-7f8eb085a7c5 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:49.024973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:49.024973Z digest=sha256:78e893cab4858c484977b8569c8eda3f937279d2ed4a72a198bbc0db1b3d13ef

Observation 2e3767e5-b00b-43a4-8dc8-8e904d53f762 · outbound

This paper cites FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:49.029983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:49.029983Z digest=sha256:ce3a91159fb649058832bdfa9e653203d2d80d64d8d4d77526784b5ad0a52736

Observation b7210c31-d1e4-47a4-91c4-8a77f36eb46b · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.601898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:49.034826Z digest=sha256:6bb627b4a5e2d3485ba38c03250822de0945dff4b50d2f97a1a8db4e85085ed2

Observation 5a9f19c5-aa01-47a0-8904-dc81364bc382 · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.580605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:49.039604Z digest=sha256:de977dbec38f6c15b94a4521840fd116256187696d56bae7dcad1a323c609ba5

Observation 8e7a93fa-78b1-47ae-93d5-681901ac99a2 · outbound

This paper cites Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:49.044638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:49.044638Z digest=sha256:2e5aee05fd83d39967a459682b21011b3071c7cc4f0d878ec09612165b7fdc09

Observation 9c1eac8c-a1ee-45a3-b851-c43abcfb45b5 · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.563376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:49.050124Z digest=sha256:7929838561b1bde5949d03ee76ef8a12a71b33bd0d02922884480e3addbfae0f

Observation 9b48ccc5-a6b8-4b00-acf9-a7f002a6b4e6 · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.544468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:49.055152Z digest=sha256:df6f1b84e93b49c9d589213563f4751b520bc3decdfdabe8b9f518c662fc1520

Observation c2a7d5fe-4728-4dda-a4cd-16d94ed44d23 · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.525570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:49.059833Z digest=sha256:8a14e6314b67a6ca9f938f357d2940c6461d95233f52159d4d26a7e3aa0d6611

Observation 1c2aa9ce-c2a9-4cdd-b070-202333293215 · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:49.064735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:49.064735Z digest=sha256:6b117d6aa3295d8ef0241cc0fdf9e6173da9b3cf8146642488bbb0b80d084f6e

Observation efe509dc-dec5-4fc3-a2d6-3bdabc967bba · outbound

This paper cites Scaling Speech-Text Pre-training with Synthetic Interleaved Data.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Scaling Speech-Text Pre-training with Synthetic Interleaved Data

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:49.070029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:49.070029Z digest=sha256:6acf3b6b7bf6786e6234776cd010cc17337ac71931e7f7c313987ae57ab2d766

Observation 129a715d-1edf-477f-a7c8-044c6d56433e · outbound

This paper cites an unresolved cited work.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-11T10:50:49.506608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T10:50:49.075609Z digest=sha256:fdb0ddca7dd488b95e0a3e0a86b8ffac9f1964ca559fa2469595c6952f9edac4

Observation b4ab5896-558d-41d0-be2a-32b1297c414f · outbound

This paper cites OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:49.080910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:49.080910Z digest=sha256:07643b789e409f05e5bb34c8d8b0b8ff00451fcf53b74742e25375b435be8e65

Observation 2d184695-04a5-4b61-bd96-95d0c1174642 · outbound

This paper cites online" 'onlinestring :=.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis online" 'onlinestring :=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:49.086842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:49.086842Z digest=sha256:2173f7496b4c5316cd101773ad8908b2f457f52cd1030f441f6c141f9f84a11e

Observation 94877811-b437-4a24-8e59-c5be6ab8d5cb · outbound

This paper cites write newline.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis write newline

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:49.092653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:49.092653Z digest=sha256:57c22212235f15c42e490e0c445906fa11ae594dee6020c2f4badf23add46673

Pith citing papers

Observation 9e006ab3-a34a-40c9-8419-e79ca9847a14 · inbound

LLM as Effective Streaming Processor: Bridging Streaming-Batch Mismatches with Group Position Encoding cites this paper.

LLM as Effective Streaming Processor: Bridging Streaming-Batch Mismatches with Group Position Encoding Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:54.401206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:54.401206Z digest=sha256:0490b704b330aaf53de03ac92df162046c2f58cb696eb8b75896c094aecdde33

Observation b18d51a4-4449-433f-b0d8-7f587de7c478 · inbound

SpeakStream: Streaming Text-to-Speech with Interleaved Data cites this paper.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.842860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.842860Z digest=sha256:3df7de5090bc4bd708701e1259c73a302aa1a3e6d2c0fed81413aea8b9077ac2

Observation 2f204fc6-c3d9-4824-bc3b-a67581bde4cc · inbound

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling cites this paper.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:52:04.503082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:02.441261Z digest=sha256:d31749f12d263073244732870191e1e8f8524f446662fce081696736b16bd809