Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:52:04.191022Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2506.12570.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:52:04.191022Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bf49560b-b22b-4a55-bc8b-a52da9d87aef · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bca8a2f3-7387-45e2-9b6f-e2b86e6e457e · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Audiolm: A language modeling approach to audio generation,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2786bbc-8f4c-4d59-89ae-3d62f6eb06f4 · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Flow matching for generative modeling,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09327ca5-7b3b-4ede-a6c2-08250ab64b36 · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Denoising diffusion probabilistic models,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b570ac72-c91a-4522-928a-e9ceca604e5f · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Neural codec language models are zero-shot text to speech synthesizers,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18b91231-a431-4427-b361-ebbc14d04113 · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling V ALL-E 2: Neural codec language models are human parity zero-shot text to speech synthesizers,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 15a4defa-8156-4b3a-902d-49cd85587af8 · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling CosyV oice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic to- kens,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b9089ddd-2bd8-46f9-94c8-a6a0b10a4ec5 · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Seed-TTS: A family of high- quality versatile speech generation models,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8bfc3e62-250f-40b5-b08d-a9fb8af1b6eb · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a4cf7c7-35da-4e8f-be7a-fe8607cfd06e · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Pseudo-autoregressive neural codec language models for efficient zero-shot text-to-speech synthesis,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 604d4ff9-0d40-492c-a316-7f8b8ab3404e · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Autoregressive speech synthesis without vector quantization,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d6f696a4-cffe-4f11-ae11-587f4073529b · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Cosyvoice 2: Scalable streaming speech synthesis with large language models,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 412ef6ee-0c2f-4a1f-b7b2-7d5c4a68b520 · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Interleaved speech-text language models are simple streaming text to speech synthesizers,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ad5cbb77-9e8e-4f73-ad1a-344020d4bdf4 · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 254d6ae7-804d-420e-bdcc-a5b0db8c718d · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Syncspeech: Low-latency and efficient dual-stream text-to-speech based on temporal masked transformer,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8d3e0bf1-4ff6-4c04-8ce3-36988ce854e1 · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Discrete audio representation as an alternative to mel-spectrograms for speaker and speech recognition,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 72d6443a-abea-45aa-87b7-dbf0dd6e7ea1 · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Autoregressive diffusion transformer for text-to-speech synthesis,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7b78f69d-f61f-4bc9-87f5-6da2c7246e5b · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Autoregressive Image Generation without Vector Quantization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67d72c9a-6b5f-4763-8b44-4b5490c36bdc · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling FELLE: autoregressive speech synthesis with token-wise coarse-to-fine flow matching,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c47da4f-c6a1-4481-95c1-2ef253ecaca8 · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Ditar: Diffusion transformer autoregressive modeling for speech generation,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36ccc01f-48dc-49d0-91d7-a45a98b8938b · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Auto-encoding variational bayes,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7082f2ce-2d21-48db-abde-1422f8f5d842 · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Tacotron: Towards end- to-end speech synthesis,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99733192-e98e-40c2-9b61-e859f987b52d · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Librispeech: an ASR corpus based on public domain audio books,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e041d1c-da61-40fa-b08a-35a0d788755a · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling V ALL-E R: robust and efficient zero-shot text-to-speech synthesis via monotonic alignment,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d419976-a6ed-4cf3-9dd7-97606372ab49 · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6a18580-b41c-42af-b442-e71d1614e80c · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling MaskGCT: Zero-shot text-to-speech with masked generative codec transformer,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7afb6a5c-1d16-46c5-9b28-d3f2c1bd98c9 · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Ramp: Retrieval-augmented mos prediction via confidence-based dynamic weighting,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 462c7a9a-7137-421b-829c-2e2aaa8cb3ec · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Ramp+: Retrieval-augmented mos prediction with prior knowledge integration,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d1701fd-bfef-49fa-abcc-e7495a5d0181 · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Conformer: Convolution-augmented Transformer for Speech Recognition
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 204c52a3-173f-43c0-a320-2e4a4bc4181b · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef7dbf33-7411-4e07-87b9-bc4450a22762 · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Robust Speech Recognition via Large-Scale Weak Supervision
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50132868-6da2-4174-9ba8-8a8d3b9a67df · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling WavLM: Large-scale self-supervised pre-training for full stack speech processing,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 63c65484-ae4f-4ce4-80a1-aa28b396b6c3 · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68ed7cca-35c7-4bf2-bf63-e61243bf22fc · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f204fc6-c3d9-4824-bc3b-a67581bde4cc · outbound
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.