Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 100 inbound Pith citation observations for arXiv:2301.02111.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T23:13:42.173985Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
20 of 20 outbound references displayed
External citation measurements
162
pith, observed 2026-08-05T02:28:24.338817Z
Observation f7c0ecb0-498a-49cb-aee6-2a874a793c51 · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers The Emotional Voices Database: Towards Controlling the Emotion Dimension in Voice Generation Systems
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80812234-f2aa-45cb-891f-e689a56cfbc5 · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers vq-wav2vec: Self-supervised learning of discrete speech representations
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af1f36ac-7d3e-43cc-bfd1-c1f8d44ad6ce · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers AudioLM: a Language Modeling Approach to Audio Generation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 45eb912f-5447-4ecb-9411-de6ec0f23d85 · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Exploring the encoding layer and loss function in end- to-end speaker and language recognition system
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd517ae2-1536-4ed0-8e25-aff5c8c62b6f · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers PaLM: Scaling Language Modeling with Pathways
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4cc583d-a536-46ed-92e7-5eaa696a5e18 · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers High Fidelity Neural Audio Compression
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05b2c44b-6903-4fa1-a6c2-0401405e2abf · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers VQTTS: high-fidelity text-to-speech synthesis with self-supervised VQ acoustic feature
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a7d60a69-6490-4353-8369-a7c041d0eb88 · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4789d3f-ed65-4a47-a9ea-7b36800571cb · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Grad-StyleSpeech: Any-speaker Adaptive Text-to-Speech Synthesis with Diffusion Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8854bd9f-f7e2-4f53-b779-cd40cc3f72a9 · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Grad-StyleSpeech: Any-speaker Adaptive Text-to-Speech Synthesis with Diffusion Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a3fe5a0c-314d-47f0-b0b5-7f48da80f4f5 · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Generative Spoken Language Modeling from Raw Audio
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84a685ac-af08-4e4c-887d-d0f151fabf4f · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Fine-grained emotion strength transfer, control and prediction for emotional speech synthesis
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a0dabbab-7521-47fb-9a03-82a0063582f2 · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Delightfultts 2: End-to-end speech synthesis with adversarial vector-quantized auto-encoders
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ac8ae89d-8bba-41b4-af8b-0fe342fc279d · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 451abf9c-4dac-461d-81da-1bc68489f169 · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8ca2d0ba-34e8-4ef2-add4-c4b84d46394e · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08d3d720-e4f8-4c04-a13d-5c1743f0d338 · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers A Survey on Neural Speech Synthesis
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eef2cd00-145a-49a9-b067-1d196dcbd143 · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Neural discrete representation learning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2ef2b2b-b19a-421a-94f9-57f323c3fd4c · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Adaspeech 4: Adaptive text to speech in zero-shot scenarios
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7a4c8b93-e642-43a0-a597-66f2f12a0c7a · outbound
Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Jingjing Xu, Xu Sun, Zhiyuan Zhang, Guangxiang Zhao, and Junyang Lin
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 433ab73b-0537-4db8-a9e7-f04b1b52da80 · inbound
Language Is Not All You Need: Aligning Perception with Language Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 40f490d2-0806-4da5-bdb6-2be2199dab83 · inbound
AudioPaLM: A Large Language Model That Can Speak and Listen Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d7b771d-c146-4787-b7af-07c3605cb699 · inbound
The Rise and Potential of Large Language Model Based Agents: A Survey Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 195
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dfb0188f-cdd3-4cdc-9bcc-16a8bbe3cb0e · inbound
DASB - Discrete Audio and Speech Benchmark Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 30f9dfec-c267-459c-b604-d7a7e2cf491e · inbound
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3d8aaba-b5f2-4ace-b0de-e7e17730300a · inbound
Moshi: a speech-text foundation model for real-time dialogue Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c039dc92-3020-4814-b481-32b6e0c061f7 · inbound
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 145
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 38eea30a-ce64-4932-b82a-c0a303b73c2d · inbound
CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0418ca03-2808-489f-aad3-f8ec4ab9c7d3 · inbound
TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-Instrument Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a34011d6-b795-4507-aa18-cf2c0f789dc4 · inbound
SparQLe: Speech Queries to Text Translation Through LLMs Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f14bc9c-e309-4a08-b023-6e708324d82e · inbound
Kimi-Audio Technical Report Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b02c52bd-2834-43fe-a9ad-050b94f8c7e7 · inbound
Perceptual implications of automatic anonymization in pathological speech Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2365d3a4-9a06-480b-a9e9-738cec03d94e · inbound
Discrete Audio Representations for Automated Audio Captioning Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b69b132-0295-4655-a468-effbb30fa8e2 · inbound
Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dd89f1e-ac5d-4aed-99c1-9262f5d8728a · inbound
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21a846d2-cdd2-4376-9909-a4cf5715e0bd · inbound
Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6fd71d8-af69-4bf9-956a-c795fc93f5f9 · inbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82899137-acd7-40a3-82ac-e034813098ca · inbound
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ce6c6bcf-db0c-4678-bf67-744e62165d24 · inbound
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb98d6b6-623a-444f-9647-735442a056ad · inbound
Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 413b2487-d923-4f77-921b-443342a17858 · inbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cc1c89c-7940-405c-8b17-9a854be9677f · inbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8688be3c-dae8-4999-810d-adf59919d480 · inbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5af4d64-86fc-478b-8394-fa81b3675b11 · inbound
OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbc6d9d5-5c85-4d49-a7c2-adb1849501d3 · inbound
EgoZero: Robot Learning from Smart Glasses Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07c2fa73-8827-4748-8982-68848deb21f1 · inbound
Voice Adaptation for Swiss German Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dca87ffb-0cc6-4a61-848e-5732005e753e · inbound
Spoken question answering for visual queries Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2a7fba6-2ea3-4d24-a852-76b0406eef20 · inbound
Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c8ee0e7-771f-415a-a6e0-7b0342511fc1 · inbound
Probing the Robustness Properties of Neural Speech Codecs Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66668e64-fd4e-4e02-80ef-9a0e7aedb0c6 · inbound
DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eedcd3b-a9f5-4164-b969-e7bb4aaf7623 · inbound
DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6ba4eca-b7b4-4add-aca3-978fb8275282 · inbound
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d902b6a-eab8-4e7b-bdbf-ab210c7bf1d1 · inbound
DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58e0589e-75b3-4209-915f-c8c78f31df42 · inbound
Zero-Shot Text-to-Speech for Vietnamese Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee2373e2-f558-4319-8718-105285ef1601 · inbound
SPBA: Utilizing Speech Large Language Model for Backdoor Attacks on Speech Classification Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 907f9a29-8e0a-4a27-99f6-cd12850217ed · inbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4705cc91-ca77-4895-aacc-1df41cdfecc3 · inbound
ViSAGe: Video-to-Spatial Audio Generation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea03eae-a2b2-4f40-8d73-c1b3035e4ac6 · inbound
Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae76af26-db9a-431e-91ec-fd57123a4bdf · inbound
Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 187d76d5-0bd2-4fa4-8b85-ed7273abebba · inbound
OpusLM: A Family of Open Unified Speech Language Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ef12dd1-6619-49f6-b981-784ba5565fb7 · inbound
Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6314508-c797-4b3c-b0c6-6201c39e2af5 · inbound
DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 294e4a4a-e8e3-4cbd-8a99-078047d1f6d8 · inbound
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 468477d5-6558-43ad-96ea-f131f8a1a28f · inbound
JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e3d7a7a-6229-458b-aee6-151703c289b9 · inbound
Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 013f9f70-da0e-40f7-ad16-7ecb4c6027bb · inbound
Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21371c99-1ebb-475b-b30e-9e9e9275736c · inbound
Traceable TTS: Toward Watermark-Free TTS with Strong Traceability Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fda98dd-0f51-48f0-b6b5-5b3af65aeb65 · inbound
SecureSpeech: Prompt-based Speaker and Content Protection Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 996fba63-73dd-41c3-8edb-b77eb667202c · inbound
Unlocking Speech Instruction Data Potential with Query Rewriting Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff6179e5-4232-4fc9-b0bf-bc23e35fc95c · inbound
ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 05b261fb-7a24-42b5-b2b3-e3be3460c77d · inbound
THAI Speech Emotion Recognition (THAI-SER) corpus Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f4ed98d-0a48-4640-9c92-26b69045f1cb · inbound
Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74a3577c-1780-4876-b29a-82e23297da18 · inbound
Autoregressive Speech Enhancement via Acoustic Tokens Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af530fdd-197d-40be-94d1-7c92f2be9d9a · inbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88e2b7df-ef30-4e6d-bbcf-ce04a4fa0eb1 · inbound
EchoVoices: Preserving Generational Voices and Memories for Seniors and Children Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39b37847-72c5-4633-94b5-1208b18484c2 · inbound
METER: Multi-modal Evidence-based Thinking and Explainable Reasoning -- Algorithm and Benchmark Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48038e2d-9ee6-464e-985b-d4c9e5d0aff9 · inbound
Step-Audio 2 Technical Report Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f0760b92-4652-409b-8b16-d7a37068a120 · inbound
Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d96c6f96-d5ae-4aa6-b75b-63251abf5781 · inbound
TTS-1 Technical Report Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95961397-4013-4c3e-9679-8b35aa91fd5a · inbound
WaveVerify: A Novel Audio Watermarking Framework for Media Authentication and Combatting Deepfakes Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3117c809-33dd-40cd-bb78-4f97974d5016 · inbound
SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f2d6d4c-ffdd-4fe9-8d08-605b21eeb050 · inbound
Adaptive Duration Model for Text Speech Alignment Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e844282-1ca3-4230-b734-949e062f358f · inbound
Next Tokens Denoising for Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8811ba4-f8e9-4233-bc25-30129c684a9b · inbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51d1501e-d1f2-45d8-8ac0-7b4eb62d7830 · inbound
Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c207100-60f2-4702-b564-334988185f9a · inbound
Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faa15446-d0e2-479a-84e9-03b0b46f337b · inbound
MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f87b3e5-0f62-484c-a2bf-6ee9e158b49b · inbound
Representing Speech Through Autoregressive Prediction of Cochlear Tokens Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7a870c0-8ae9-4cb2-b7d2-6361397113fe · inbound
CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92059eaf-e1d3-4fd6-a32d-58e5de8d0d8c · inbound
CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caa4fed2-b6f6-4b3a-bfee-25ed8bfb1ed6 · inbound
FreeTalk:A plug-and-play and black-box defense against speech synthesis attacks Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 801c9947-2627-41c8-acd0-93c70e48918f · inbound
MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d9d3ebc-8870-446b-8f2e-5308dc2057af · inbound
Enhancing Speech Large Language Models through Reinforced Behavior Alignment Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 896d6676-1b94-453e-be33-5b385a527b2b · inbound
WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 451c10cf-ad20-40b1-81ef-1a87c8c042a6 · inbound
Cloning a Conversational Voice AI Agent from Call\,Recording Datasets for Telesales Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d66cf36-f702-433f-8c07-6879722d55a0 · inbound
Effectively obtaining acoustic, visual and textual data from videos Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 114
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f489a456-69c3-44f8-9163-9a4a535cf71a · inbound
DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96f8f981-3f28-4c28-b8f2-5f1b8e04085a · inbound
Testing chatbots on the creation of encoders for audio conditioned image generation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 104
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb4c225b-9529-45bc-b444-56a161c55a5a · inbound
DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb2213ca-41f2-4def-9e72-380d921a6b27 · inbound
Length-Aware Rotary Position Embedding for Text-Speech Alignment Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e51b81a-7193-40de-9309-2975f587cb7b · inbound
CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2de345b8-464a-42fa-ad68-22a579d80986 · inbound
CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b3821e8-c4cb-4589-b362-5bfc4485c04c · inbound
UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7e93bc3-2ecb-4ced-b1d1-41c14daf4228 · inbound
Speculative Coupled Decoding for Training-Free Lossless Acceleration of Autoregressive Visual Generation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9289ac55-4351-41a5-85b6-75094684eb58 · inbound
Two-Dimensional Quantization for Geometry-Aware Audio Coding Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 401bb360-ae68-4c51-844e-693ef14f636f · inbound
Aliasing-Free Neural Audio Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c0dad54a-e6e0-4ac5-b440-8a7c80231c7f · inbound
Aliasing-Free Neural Audio Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 407faa63-45d7-4036-a722-81c0e57e3ba1 · inbound
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07662bee-7395-480a-923e-d7af4d78f88c · inbound
On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 237646c5-2a90-4e69-977e-307008aef4ea · inbound
Qwen3-TTS Technical Report Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e58362e4-73da-4412-a91f-d82f8a5896f2 · inbound
SKETCH: Semantic Key-Point Conditioning for Long-Horizon Vessel Trajectory Prediction Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e3f6799-ecbc-415d-ab95-7fdd18c8b98a · inbound
AUHead: Realistic Emotional Talking Head Generation via Action Units Control Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc76f225-cba5-4588-b0a1-fb50438b7dc1 · inbound
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cd1e6aa-b77f-48ac-8e1a-832191215d75 · inbound
Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3808e4c4-f71f-4a9c-8fba-5651c4848a1c · inbound
Borderless Long Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59891d00-1594-4f9d-ba16-054ea3478155 · inbound
Voxtral TTS Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 04b87d48-ae9f-4f0b-b171-95f4075104c7 · inbound
Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8a133291-d5de-4349-9bcc-0ea6f55e40ce · inbound
WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 727feb1c-1375-4125-b8b5-28409fa22d77 · inbound
HAFM: Hierarchical Autoregressive Foundation Model for Music Accompaniment Generation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51985833-c37d-4cbb-bb7b-9174ec1da7ca · inbound
ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.