Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T10:50:49.092653Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 3 inbound Pith citation observations for arXiv:2412.16102.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T10:50:49.092653Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:56:54.401206Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T00:52:04.433992Z
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5abad7c9-a9d2-4dfe-a222-0e9041616592 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1fe770c-7af2-4e8f-85e9-86fa99ccdc68 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 608e2db3-527c-4222-8601-e0d3c92feb6b · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 68bac971-2d16-4aa3-b4b2-16674cbe49a3 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a85879f1-7564-47f2-8945-97aa42eeac38 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4148da6-3c4e-4d9f-bdbd-5b7ac08b07af · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c953c41b-7e6d-4422-8d8f-dbe3ef932164 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d7e6913d-f226-4a9e-9dc8-f132b54e1678 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a27e6bf5-dd16-48a4-b502-0299cba8771c · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3f5245a7-f4a1-49f4-8297-a9b7af24af27 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99cc4685-422a-4cd0-91c2-c008420c40ce · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2b146d94-fd67-4d8a-a121-1cba0b0853a8 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49863daa-bcf8-47f1-b54b-0375eb04efe1 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Assael, Brendan Shillingford, and 1 others
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a3be91f4-229e-4d20-bd13-2aea0368aa51 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa140664-ae2b-4f37-9c28-9570a3db10ad · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Zero-Shot Text-to-Speech from Continuous Text Streams
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 904b4ade-b7cf-4e5e-a578-f8fe0250750e · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 90863e13-c9b8-4101-861d-0d5762bdc042 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e9e6cbd4-2ac1-4c1a-99b9-6f5ca1aa60a4 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5e1f0168-0105-408f-bc37-96ad5c19245a · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73d3a66c-a74d-4113-a959-0ae78c8ec21d · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8163073-2598-4678-8751-2845f62ee83f · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 76646904-ff98-4c32-b4a0-1f8f9001f0af · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a09c2152-1153-4bc3-b51c-7e98aaefb144 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3cefdc99-1888-47a2-8bfa-07772a00ec2c · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f526b32d-ca66-4533-9cf7-164a3f0e8e8d · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4db010cb-b496-40ba-8481-68349392e9c0 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Weiss, and 1 others
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 95777c88-558a-42a0-ad13-1da4a7653408 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b9c2baac-c41f-4619-89d2-7fd1ab8300bf · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 68c1d559-5220-4df1-83cd-6310dbe9cfbe · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b8f29230-c0fb-4d4c-935f-968d993dab08 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3c35198d-3d1b-4137-8afa-9f31ce5e0121 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0700a401-29e4-4052-8d74-92c5d4a6da43 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2020cdbb-4002-4a58-8b59-cb0e193aea04 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30cf56df-974d-45e8-a5af-eb694d606638 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2f843f18-b863-40f1-abeb-60f527d6c75b · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8b799533-3325-4bde-8e14-8d9bdcf774ef · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 02341cbf-8943-44cd-bf2d-b7c3eb318834 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41dacecb-7a32-48e5-bd4a-bd2f39fe52eb · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Autoregressive Speech Synthesis without Vector Quantization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9ab39a5-cec5-4a91-aef3-9d769a3a6eb1 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Spirit LM: Interleaved Spoken and Written Language Model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbd5eb9f-a4ac-4791-babe-43352f2a4842 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e4210081-0b81-4d23-ab2d-7120bbee6044 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d01f23f2-e972-4f4b-b657-dbbc12cc4cce · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 827437ce-fad0-4902-b85c-90aa0079c1a8 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Weiss, and 1 others
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 06bef987-0894-470b-9668-add4ed54f302 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9d98538-0be9-4d7b-bbdb-f1635063d93a · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0447ea0a-cbd9-4840-9ce2-69080204b639 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1a01b47c-a422-473e-ad87-7f8eb085a7c5 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e3767e5-b00b-43a4-8dc8-8e904d53f762 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7210c31-d1e4-47a4-91c4-8a77f36eb46b · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5a9f19c5-aa01-47a0-8904-dc81364bc382 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8e7a93fa-78b1-47ae-93d5-681901ac99a2 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c1eac8c-a1ee-45a3-b851-c43abcfb45b5 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9b48ccc5-a6b8-4b00-acf9-a7f002a6b4e6 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c2a7d5fe-4728-4dda-a4cd-16d94ed44d23 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1c2aa9ce-c2a9-4cdd-b070-202333293215 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efe509dc-dec5-4fc3-a2d6-3bdabc967bba · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Scaling Speech-Text Pre-training with Synthetic Interleaved Data
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 129a715d-1edf-477f-a7c8-044c6d56433e · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b4ab5896-558d-41d0-be2a-32b1297c414f · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d184695-04a5-4b61-bd96-95d0c1174642 · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis online" 'onlinestring :=
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94877811-b437-4a24-8e59-c5be6ab8d5cb · outbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis write newline
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e006ab3-a34a-40c9-8419-e79ca9847a14 · inbound
LLM as Effective Streaming Processor: Bridging Streaming-Batch Mismatches with Group Position Encoding Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b18d51a4-4449-433f-b0d8-7f587de7c478 · inbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f204fc6-c3d9-4824-bc3b-a67581bde4cc · inbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.