Pith. sign in

Paper Citation Record · LEDGER

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

As of 8 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 100 inbound Pith citation observations for arXiv:2301.02111.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2301.02111 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 120 of 120 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 100 of 191 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:13:42.173985Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact9
  • verified fuzzy8
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

162
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation f7c0ecb0-498a-49cb-aee6-2a874a793c51 · outbound

This paper cites The Emotional Voices Database: Towards Controlling the Emotion Dimension in Voice Generation Systems.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers The Emotional Voices Database: Towards Controlling the Emotion Dimension in Voice Generation Systems

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:31:07.124381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:09a29828a6a4cd4bee24c460aa7b1069de3dbc4146167e21f7f132da83eec257

Observation 80812234-f2aa-45cb-891f-e689a56cfbc5 · outbound

This paper cites vq-wav2vec: Self-supervised learning of discrete speech representations.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers vq-wav2vec: Self-supervised learning of discrete speech representations

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:31:07.157127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:deb30ae0e0e1229e926c9c9f5dea4fb514faa2dc6b0cfd41998b03f9132fa3b6

Observation af1f36ac-7d3e-43cc-bfd1-c1f8d44ad6ce · outbound

This paper cites AudioLM: a Language Modeling Approach to Audio Generation.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers AudioLM: a Language Modeling Approach to Audio Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:31:07.104833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:e885b07a903492eae39b105fe93fc677d22589d46edfd0fb59bbc6c6573329c4

Observation 45eb912f-5447-4ecb-9411-de6ec0f23d85 · outbound

This paper cites Exploring the encoding layer and loss function in end- to-end speaker and language recognition system.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Exploring the encoding layer and loss function in end- to-end speaker and language recognition system

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:31:07.144836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:1309a921bed884c3caa32ede27a9b419a78e354298146e00ebce00cf02e30d22

Observation dd517ae2-1536-4ed0-8e25-aff5c8c62b6f · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers PaLM: Scaling Language Modeling with Pathways

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:31:07.111554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:1aa19cd9834b28340fd4a8f471aa9d7b142e1467d7ed1fb8b6ddbbefcb020a95

Observation b4cc583d-a536-46ed-92e7-5eaa696a5e18 · outbound

This paper cites High Fidelity Neural Audio Compression.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers High Fidelity Neural Audio Compression

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:49:52.306785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:818a5809f70eb88487687af2e87aad5bc07fa5596ac787a695f348a534454656

Observation 05b2c44b-6903-4fa1-a6c2-0401405e2abf · outbound

This paper cites VQTTS: high-fidelity text-to-speech synthesis with self-supervised VQ acoustic feature.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers VQTTS: high-fidelity text-to-speech synthesis with self-supervised VQ acoustic feature

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:31:07.166080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:48834c4c78b4c12aeb381eb5888b8c4886a6ac5f5c4fc1403c258ef23f197134

Observation a7d60a69-6490-4353-8369-a7c041d0eb88 · outbound

This paper cites Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai, Kushal Lakhotia, Ruslan Salakhutdinov, and Abdelrahman Mohamed

Reference 8

Resolution
verified exact
doi, observed 2026-05-13T01:31:07.078073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:0e500b6fce3b59895225cf8487791036baabc52962c8554192c1b369486898f3

Observation d4789d3f-ed65-4a47-a9ea-7b36800571cb · outbound

This paper cites Grad-StyleSpeech: Any-speaker Adaptive Text-to-Speech Synthesis with Diffusion Models.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Grad-StyleSpeech: Any-speaker Adaptive Text-to-Speech Synthesis with Diffusion Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:31:07.090934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:dd51476aff71f29770d650261b16ab2c61e6d3e3ba37800ecfc2cdc14505e576

Observation 8854bd9f-f7e2-4f53-b779-cd40cc3f72a9 · outbound

This paper cites Grad-StyleSpeech: Any-speaker Adaptive Text-to-Speech Synthesis with Diffusion Models.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Grad-StyleSpeech: Any-speaker Adaptive Text-to-Speech Synthesis with Diffusion Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:31:07.069766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:65e01a7fd3d8c3338e4e52a9d07dc77b4c92583d73bddfab96bd1112acdf2553

Observation a3fe5a0c-314d-47f0-b0b5-7f48da80f4f5 · outbound

This paper cites Generative Spoken Language Modeling from Raw Audio.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Generative Spoken Language Modeling from Raw Audio

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:31:07.097756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:1dae429294ba2007cf9325979af8004fea6f90a7d028c18f9a5b6016f7ae505f

Observation 84a685ac-af08-4e4c-887d-d0f151fabf4f · outbound

This paper cites Fine-grained emotion strength transfer, control and prediction for emotional speech synthesis.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Fine-grained emotion strength transfer, control and prediction for emotional speech synthesis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:31:07.134672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:475a1fb067df9229e53f77428206f8caa66b75cd2edd9ae22289adac963be799

Observation a0dabbab-7521-47fb-9a03-82a0063582f2 · outbound

This paper cites Delightfultts 2: End-to-end speech synthesis with adversarial vector-quantized auto-encoders.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Delightfultts 2: End-to-end speech synthesis with adversarial vector-quantized auto-encoders

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:31:07.139596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:c50058ee58f5638011a8a7a733de3de01e2ec8adf7392a8530ef02eff5637e0b

Observation ac8ae89d-8bba-41b4-af8b-0fe342fc279d · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T01:31:07.062215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:2e1d2e726177a4bab9c8c056308a67487be901750733ff11cfe1a49087316c06

Observation 451abf9c-4dac-461d-81da-1bc68489f169 · outbound

This paper cites an unresolved cited work.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-13T01:31:07.148938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:0c80660e254c77d6a43174010dbef3dd65d64042cfb79ba58e76b90ce99f2909

Observation 8ca2d0ba-34e8-4ef2-add4-c4b84d46394e · outbound

This paper cites Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Daniel Povey, Arnab Ghoshal, Gilles Boulianne, Lukas Burget, Ondrej Glembek, Nagendra Goel, Mirko Hannemann, Petr Motlicek, Yanmin Qian, Petr Schwarz, et al

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:31:07.152932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:d896fc5acb264232e417e3a4ac4d74a9b79c7cde24b2920d9d96c352bccd9344

Observation 08d3d720-e4f8-4c04-a13d-5c1743f0d338 · outbound

This paper cites A Survey on Neural Speech Synthesis.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers A Survey on Neural Speech Synthesis

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:31:07.130156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:6c46aa1090d8743de3247202f738871ba6984b117d39d99e75c7623c26334f91

Observation eef2cd00-145a-49a9-b067-1d196dcbd143 · outbound

This paper cites Neural discrete representation learning.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Neural discrete representation learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:31:07.161342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:ba2a3b56a89c9a028bfad9dac073d060bf7247d76fccd0f320c98f2ec0cae0aa

Observation e2ef2b2b-b19a-421a-94f9-57f323c3fd4c · outbound

This paper cites Adaspeech 4: Adaptive text to speech in zero-shot scenarios.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Adaspeech 4: Adaptive text to speech in zero-shot scenarios

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T01:31:07.170645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:96de30ffcd8eb12825de4ff6d4f596f5adea563c473ef50a26926ef4234a834a

Observation 7a4c8b93-e642-43a0-a597-66f2f12a0c7a · outbound

This paper cites Jingjing Xu, Xu Sun, Zhiyuan Zhang, Guangxiang Zhao, and Junyang Lin.

Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers Jingjing Xu, Xu Sun, Zhiyuan Zhang, Guangxiang Zhao, and Junyang Lin

Reference 20

Resolution
verified exact
doi, observed 2026-05-13T01:31:07.082725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:31:06.635613Z digest=sha256:7fbcbd5a052cebc642fd365292dc729747a06f62dd5cd9c4bff0f222d23787c4

Pith citing papers

Observation 433ab73b-0537-4db8-a9e7-f04b1b52da80 · inbound

Language Is Not All You Need: Aligning Perception with Language Models cites this paper.

Language Is Not All You Need: Aligning Perception with Language Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:32:22.867191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T18:32:22.813668Z digest=sha256:8c54f7c575641cfa1a9134ec5dc6f4426dde06e5470bd19d8c0db3644cfecca0

Observation 40f490d2-0806-4da5-bdb6-2be2199dab83 · inbound

AudioPaLM: A Large Language Model That Can Speak and Listen cites this paper.

AudioPaLM: A Large Language Model That Can Speak and Listen Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:07:57.904544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T07:07:57.800866Z digest=sha256:054df158b359b038f2140ef5a43ccf7edaa6d4700f1792959d727a62d4aae89d

Observation 1d7b771d-c146-4787-b7af-07c3605cb699 · inbound

The Rise and Potential of Large Language Model Based Agents: A Survey cites this paper.

The Rise and Potential of Large Language Model Based Agents: A Survey Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 195

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:31:07.171839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T10:47:44.152066Z digest=sha256:316f368630024a327335ff8b7fcf0e687626a7e183e8c74fc98bbccd822bbd5d

Observation dfb0188f-cdd3-4cdc-9bcc-16a8bbe3cb0e · inbound

DASB - Discrete Audio and Speech Benchmark cites this paper.

DASB - Discrete Audio and Speech Benchmark Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-24T00:28:39.486547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:7ab8d9a6790c7b0d45634d881cd7b6d6df97a5e7640a08f87e05d09be2ebf7a1

Observation 30f9dfec-c267-459c-b604-d7a7e2cf491e · inbound

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS cites this paper.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:33:25.623334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:c16948ea671b1c2acf3e6122e6a6f134ccf62e5a56251060297de05570f23cc6

Observation d3d8aaba-b5f2-4ace-b0de-e7e17730300a · inbound

Moshi: a speech-text foundation model for real-time dialogue cites this paper.

Moshi: a speech-text foundation model for real-time dialogue Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:31:07.171839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T08:13:21.962488Z digest=sha256:5204abe6223137d0d7272816216ab99f8ba32b44cecb67d1dd35b7ca9c48f475

Observation c039dc92-3020-4814-b481-32b6e0c061f7 · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 145

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:06:41.538849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:15172ff65aa2402cc0e75dbaf639d19f058672b7202975556d694800aa9edd49

Observation 38eea30a-ce64-4932-b82a-c0a303b73c2d · inbound

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models cites this paper.

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:19:09.534650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T06:19:09.507440Z digest=sha256:7dc8e4b678e4a5733bde6a0b61a3d9801a732ed6c7b2499871366eaeb6a6f9a2

Observation 0418ca03-2808-489f-aad3-f8ec4ab9c7d3 · inbound

TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-Instrument cites this paper.

TokenSynth: A Token-based Neural Synthesizer for Instrument Cloning and Text-to-Instrument Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T23:13:42.173985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:13:42.173985Z digest=sha256:2235f373bc9e30ecba6c7895c45f6ffe5dca26d2bc2659ed6aa009977a1a2b3c

Observation a34011d6-b795-4507-aa18-cf2c0f789dc4 · inbound

SparQLe: Speech Queries to Text Translation Through LLMs cites this paper.

SparQLe: Speech Queries to Text Translation Through LLMs Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T22:05:23.870284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:05:23.870284Z digest=sha256:b05ea3eee5f5bdc68cf8375a16cc9aca02db6ce0998fb04a60738f688636a72a

Observation 6f14bc9c-e309-4a08-b023-6e708324d82e · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:31:07.171839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:25c031c37a657aecd440d809cf40b5864af23fdceadbe87a215558e65615fd19

Observation b02c52bd-2834-43fe-a9ad-050b94f8c7e7 · inbound

Perceptual implications of automatic anonymization in pathological speech cites this paper.

Perceptual implications of automatic anonymization in pathological speech Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:01:54.549853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T17:58:15.380780Z digest=sha256:9727fe1ae9f822c29ccaacaaf8aaac32b88d76ecee254152d63c1980435a69e9

Observation 2365d3a4-9a06-480b-a9e9-738cec03d94e · inbound

Discrete Audio Representations for Automated Audio Captioning cites this paper.

Discrete Audio Representations for Automated Audio Captioning Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:39.691319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:39.691319Z digest=sha256:17d76adf1895e67b53cbbe0f7ef04c1710783f19861ee460977cf58e27aa3fd4

Observation 9b69b132-0295-4655-a468-effbb30fa8e2 · inbound

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding cites this paper.

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:15.575777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:15.575777Z digest=sha256:160f5f06c783990b6f1e58c865b00a688c79569b251d72a8a594ab6d2591c706

Observation 9dd89f1e-ac5d-4aed-99c1-9262f5d8728a · inbound

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition cites this paper.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.998249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.998249Z digest=sha256:539d7fc9541acabb07b25c973246d048320d05bc189e14ed1694b47fa59e8784

Observation 21a846d2-cdd2-4376-9909-a4cf5715e0bd · inbound

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English cites this paper.

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:43:04.468219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:43:04.468219Z digest=sha256:96d23b2fb34fb76c745a843667fc0b8bdf0c46e4ce06909bdddd8dcc81aa97c2

Observation f6fd71d8-af69-4bf9-956a-c795fc93f5f9 · inbound

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information cites this paper.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.522496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.522496Z digest=sha256:514f1163c0f52877e2a262f74383dcafd8e25c510cc2099ec23219c55a08d2bb

Observation 82899137-acd7-40a3-82ac-e034813098ca · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:27:25.468581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:f41fbf97c898939a4eb3fbd79d20f14f498302a621ca94fdd981b02fea88f8e5

Observation ce6c6bcf-db0c-4678-bf67-744e62165d24 · inbound

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt cites this paper.

MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:38.473082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:38.473082Z digest=sha256:aad3c68505ceaaf29a01ad6195fc3b3c921d9994ef604aec4b9e7a3c8cbf3707

Observation fb98d6b6-623a-444f-9647-735442a056ad · inbound

Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis cites this paper.

Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:26:03.465837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:26:03.465837Z digest=sha256:abe7fbbf1dc06b67318297c4a169187bbf6acc8ad2298b71cec175b1a207867a

Observation 413b2487-d923-4f77-921b-443342a17858 · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:50.254673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:50.254673Z digest=sha256:249c7dffae196870cf123fe17ab7a4bff4df4ae893b8471c7f769511034848ab

Observation 6cc1c89c-7940-405c-8b17-9a854be9677f · inbound

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment cites this paper.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:53.471613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:53.471613Z digest=sha256:465fd8c8c502214b5d73431ce25406bbfed53bc3f06e86f08e1d6977e32191b7

Observation 8688be3c-dae8-4999-810d-adf59919d480 · inbound

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling cites this paper.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.172894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.172894Z digest=sha256:f0fbcc41e67042f40c09053a3ba47c3f6bbc2dd6fefd176b3d3633c96e37bf3e

Observation c5af4d64-86fc-478b-8394-fa81b3675b11 · inbound

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction cites this paper.

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:32.507140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:01:32.507140Z digest=sha256:94a8eeabd67a7c8b0c8565e3e856b45d565d54bbfb6bfe8ab24948a18721c436

Observation dbc6d9d5-5c85-4d49-a7c2-adb1849501d3 · inbound

EgoZero: Robot Learning from Smart Glasses cites this paper.

EgoZero: Robot Learning from Smart Glasses Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:00:57.267762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:00:57.267762Z digest=sha256:39a56e0106a7983ce3935f290d5f35238119a85d3ca397081f48950eeb783d0c

Observation 07c2fa73-8827-4748-8982-68848deb21f1 · inbound

Voice Adaptation for Swiss German cites this paper.

Voice Adaptation for Swiss German Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:04.803829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:04.803829Z digest=sha256:1eb248b519ee785528e3ef8ed89c7bbccf1b3127fec027d8c902e1a8c81f8093

Observation dca87ffb-0cc6-4a61-848e-5732005e753e · inbound

Spoken question answering for visual queries cites this paper.

Spoken question answering for visual queries Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:39.261148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:39.261148Z digest=sha256:04cda9b6ac2b9591a4c406d52d52d1b402407e04f6cb900c881e79b947c431ba

Observation d2a7fba6-2ea3-4d24-a852-76b0406eef20 · inbound

Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes cites this paper.

Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:29.146399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:29.146399Z digest=sha256:d0ff2d424d4a30740834ea578a4cf27f5fd32c72aea27a87efa88cde26e02010

Observation 0c8ee0e7-771f-415a-a6e0-7b0342511fc1 · inbound

Probing the Robustness Properties of Neural Speech Codecs cites this paper.

Probing the Robustness Properties of Neural Speech Codecs Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:31:37.116039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:31:37.116039Z digest=sha256:bd844dcfcef0ae5e001338dfc8b593c885e050fb8b0c7e6239af143072b6b01e

Observation 66668e64-fd4e-4e02-80ef-9a0e7aedb0c6 · inbound

DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec cites this paper.

DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:39.170430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:39.170430Z digest=sha256:facfc60f8232ed6f4d254f2675dbe3c3453d9f5c9c63800d32c7415aedb58d78

Observation 5eedcd3b-a9f5-4164-b969-e7bb4aaf7623 · inbound

DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model cites this paper.

DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:56.274668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:56.274668Z digest=sha256:009f16bf4630157b0608fef53b09310d790bec861e9a26cc8ea7653dc339c8d4

Observation f6ba4eca-b7b4-4add-aca3-978fb8275282 · inbound

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching cites this paper.

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:57.777008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:57.777008Z digest=sha256:57bb983f69723c3a8b7601788a4b4a78827faab5831f5f650b0a5065d56a451d

Observation 7d902b6a-eab8-4e7b-bdbf-ab210c7bf1d1 · inbound

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation cites this paper.

DS-TTS: Zero-Shot Speaker Style Adaptation from Voice Clips via Dynamic Dual-Style Feature Modulation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:29.155283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:29.155283Z digest=sha256:2ddea6acda0138d49f2726b79b1dd6d887bc9593368cf70a5d6a4997e9349a4f

Observation 58e0589e-75b3-4209-915f-c8c78f31df42 · inbound

Zero-Shot Text-to-Speech for Vietnamese cites this paper.

Zero-Shot Text-to-Speech for Vietnamese Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:05.229980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:50:05.229980Z digest=sha256:bd2ec7826d6059a6ab81aab1f7d38a327aba8eebb2ac51d3e4b4a6ce6b09c166

Observation ee2373e2-f558-4319-8718-105285ef1601 · inbound

SPBA: Utilizing Speech Large Language Model for Backdoor Attacks on Speech Classification Models cites this paper.

SPBA: Utilizing Speech Large Language Model for Backdoor Attacks on Speech Classification Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:18:04.441044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:18:04.441044Z digest=sha256:6c7c7e1af6b000f18bcfba0518215e1e60f3e51cb9aab66855fcab83a4a586e5

Observation 907f9a29-8e0a-4a27-99f6-cd12850217ed · inbound

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model cites this paper.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.171753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.171753Z digest=sha256:ce77882db2c723dc8d187a3e88c8cd76b99dbecf04f591bf31c1bf24c24f0664

Observation 4705cc91-ca77-4895-aacc-1df41cdfecc3 · inbound

ViSAGe: Video-to-Spatial Audio Generation cites this paper.

ViSAGe: Video-to-Spatial Audio Generation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T01:04:53.342898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:04:53.342898Z digest=sha256:506cd973d62512dc34e610671cd5ca6a41ba3799d369055d310c924cc8d7adc9

Observation 8ea03eae-a2b2-4f40-8d73-c1b3035e4ac6 · inbound

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification cites this paper.

Investigation of Zero-shot Text-to-Speech Models for Enhancing Short-Utterance Speaker Verification Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:38.647520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:38.647520Z digest=sha256:755d922a141b0bd55b53929e307dbfc7bce7e7967ccebf3909425c5f925ac02d

Observation ae76af26-db9a-431e-91ec-fd57123a4bdf · inbound

Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition cites this paper.

Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:33.903012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:33.903012Z digest=sha256:525f75b06c141eafbf1b5e71a3e2c2d2c78161ea84e4c4bd52e22cd2b6abf5de

Observation 187d76d5-0bd2-4fa4-8b85-ed7273abebba · inbound

OpusLM: A Family of Open Unified Speech Language Models cites this paper.

OpusLM: A Family of Open Unified Speech Language Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:37.245740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:37.245740Z digest=sha256:b608ec0b16c92acaedc22204b4b7d5699465c7d9cca671b93bf34ddcf65f67eb

Observation 8ef12dd1-6619-49f6-b981-784ba5565fb7 · inbound

Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation cites this paper.

Vo-Ve: An Explainable Voice-Vector for Speaker Identity Evaluation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:40.868064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:40.868064Z digest=sha256:e589f6a1d49dee035f734a5f9e78b30c40aa132de19dbb1cbe2ca7005f85ab58

Observation f6314508-c797-4b3c-b0c6-6201c39e2af5 · inbound

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding cites this paper.

DiffSoundStream: Efficient Speech Tokenization via Diffusion Decoding Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:11:58.112443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:11:58.112443Z digest=sha256:ffbea8e8084729cea85381e5198f2e8a0470fd6b84d6620120b00c9150d5698a

Observation 294e4a4a-e8e3-4cbd-8a99-078047d1f6d8 · inbound

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs cites this paper.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:36.176279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:36.176279Z digest=sha256:380aed0d12aa19ac604e68122b413f6a5c53969849df54794f51144f13476193

Observation 468477d5-6558-43ad-96ea-f131f8a1a28f · inbound

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching cites this paper.

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-22T00:50:51.058050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T00:46:39.196042Z digest=sha256:32a045d64b9bb3d98ccaf3c63681f00165bcdac557560b0d81da5009c0252928

Observation 1e3d7a7a-6229-458b-aee6-151703c289b9 · inbound

Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis cites this paper.

Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:25:26.955694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:25:26.955694Z digest=sha256:a45c5d28bd011831120cff436664c75bdee229717c74dabc512beeaa67f907d1

Observation 013f9f70-da0e-40f7-ad16-7ecb4c6027bb · inbound

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges cites this paper.

Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:48.037703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:48.037703Z digest=sha256:d1d2d0479704361c964ea7fe30773c866d0f9ca71a72b0d2c2d90819657676e1

Observation 21371c99-1ebb-475b-b30e-9e9e9275736c · inbound

Traceable TTS: Toward Watermark-Free TTS with Strong Traceability cites this paper.

Traceable TTS: Toward Watermark-Free TTS with Strong Traceability Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:06:18.853667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:06:18.853667Z digest=sha256:c8673aefdfa1875a2180ed079bcef9df30c0d87a105a75df6f1722c69df4a0f3

Observation 6fda98dd-0f51-48f0-b6b5-5b3af65aeb65 · inbound

SecureSpeech: Prompt-based Speaker and Content Protection cites this paper.

SecureSpeech: Prompt-based Speaker and Content Protection Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:36:07.588171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:36:07.588171Z digest=sha256:303c298c6fbf6f5ea5e5649c9761323756e45c583a9779d9c0ec8ad9aa887865

Observation 996fba63-73dd-41c3-8edb-b77eb667202c · inbound

Unlocking Speech Instruction Data Potential with Query Rewriting cites this paper.

Unlocking Speech Instruction Data Potential with Query Rewriting Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:21:35.186108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:21:35.186108Z digest=sha256:a6c65b5767653597db7b3c84925fd1e6b0d89695de848434432f1e4634f01ddc

Observation ff6179e5-4232-4fc9-b0bf-bc23e35fc95c · inbound

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching cites this paper.

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-19T04:32:03.689254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T04:29:41.285194Z digest=sha256:3d2bbc05cb7932a4742739cdf800067211f057e4de87392f178b6028536f51a5

Observation 05b261fb-7a24-42b5-b2b3-e3be3460c77d · inbound

THAI Speech Emotion Recognition (THAI-SER) corpus cites this paper.

THAI Speech Emotion Recognition (THAI-SER) corpus Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T17:57:51.209447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:57:51.209447Z digest=sha256:a7a1a4cadd1d7fb31482d8040139f1923c96978a83f0114ac951cc36b759f761

Observation 8f4ed98d-0a48-4640-9c92-26b69045f1cb · inbound

Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine cites this paper.

Task-Specific Audio Coding for Machines: Machine-Learned Latent Features Are Codes for That Machine Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:47:34.327894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:47:34.327894Z digest=sha256:8e2d99223a04802bba8da42cc78d8f06d9fed70e00211f0faffa254eee11bb50

Observation 74a3577c-1780-4876-b29a-82e23297da18 · inbound

Autoregressive Speech Enhancement via Acoustic Tokens cites this paper.

Autoregressive Speech Enhancement via Acoustic Tokens Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:21.629566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:21.629566Z digest=sha256:77adbedc1551bdb0fb40697374195ee46f132efd60b2e01439b8bf28f373710a

Observation af530fdd-197d-40be-94d1-7c92f2be9d9a · inbound

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis cites this paper.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.353019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.353019Z digest=sha256:97da65d5b090a50cbd15bdcd1cacd3723a314ba3475e6a7bfdbc1ce48537aea9

Observation 88e2b7df-ef30-4e6d-bbcf-ce04a4fa0eb1 · inbound

EchoVoices: Preserving Generational Voices and Memories for Seniors and Children cites this paper.

EchoVoices: Preserving Generational Voices and Memories for Seniors and Children Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:42:20.987345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:42:20.987345Z digest=sha256:43f8860ca4e46b0273e4d01003a4c5cbbd2542e2ee94a0daf34059f269c0d287

Observation 39b37847-72c5-4633-94b5-1208b18484c2 · inbound

METER: Multi-modal Evidence-based Thinking and Explainable Reasoning -- Algorithm and Benchmark cites this paper.

METER: Multi-modal Evidence-based Thinking and Explainable Reasoning -- Algorithm and Benchmark Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:18:08.572945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:18:08.572945Z digest=sha256:71142a12d7f3fcfb2b5a207d1528ddcc824b62be6466c0095615ff8f7672e16e

Observation 48038e2d-9ee6-464e-985b-d4c9e5d0aff9 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:59:51.056488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:b1d38498bf051e4fd69b227e3981a06b129d9c29d651a14cca21e2a0b6a15b77

Observation f0760b92-4652-409b-8b16-d7a37068a120 · inbound

Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages cites this paper.

Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:31.452735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:14:31.452735Z digest=sha256:c0926f01f74f1067ae45db75505b64eb5a49c093ce231e84fde9fe682ca7847d

Observation d96c6f96-d5ae-4aa6-b75b-63251abf5781 · inbound

TTS-1 Technical Report cites this paper.

TTS-1 Technical Report Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.886470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.886470Z digest=sha256:0d02622237a5352917dbcb5cef956218537e7258068a96ee61103748c9dc7753

Observation 95961397-4013-4c3e-9679-8b35aa91fd5a · inbound

WaveVerify: A Novel Audio Watermarking Framework for Media Authentication and Combatting Deepfakes cites this paper.

WaveVerify: A Novel Audio Watermarking Framework for Media Authentication and Combatting Deepfakes Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:38.051881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:38.051881Z digest=sha256:be2cfaa4017531b720b90007e5c0d13a57a399493a787674b9994ed4a7b4f4cd

Observation 3117c809-33dd-40cd-bb78-4f97974d5016 · inbound

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods cites this paper.

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T12:49:22.407513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:49:22.407513Z digest=sha256:c762fa7674cacf64626be5617aac4936d9fa7962b38d9601eccfa024966dc51a

Observation 7f2d6d4c-ffdd-4fe9-8d08-605b21eeb050 · inbound

Adaptive Duration Model for Text Speech Alignment cites this paper.

Adaptive Duration Model for Text Speech Alignment Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:09.160142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:09.160142Z digest=sha256:08db225693fe46bed26929f1326fccb3d54c68b11807124aeea7bca562ff982c

Observation 2e844282-1ca3-4230-b734-949e062f358f · inbound

Next Tokens Denoising for Speech Synthesis cites this paper.

Next Tokens Denoising for Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:27.428195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:27.428195Z digest=sha256:2592aa80068a9dea41ecc7a3d3b93e647de7e37a0804f1a7231d8ff02dd5e0b8

Observation a8811ba4-f8e9-4233-bc25-30129c684a9b · inbound

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation cites this paper.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.948925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.948925Z digest=sha256:cd107841befdfcaee40b1e54632e78b5ea87aeef4da87661e6e676ea224f9cdd

Observation 51d1501e-d1f2-45d8-8ac0-7b4eb62d7830 · inbound

Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation cites this paper.

Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T05:40:56.949098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:40:56.949098Z digest=sha256:69877aec4f69fbbb8c10b04ffd2561290e0775ba7cc242032083a4a58a3869d5

Observation 5c207100-60f2-4702-b564-334988185f9a · inbound

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation cites this paper.

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T23:43:05.255488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:43:05.255488Z digest=sha256:a85a21270c2b43bea1ec09ee293e424d18277f41452afcd78f8cc334db75ebff

Observation faa15446-d0e2-479a-84e9-03b0b46f337b · inbound

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts cites this paper.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:50.064063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:50.064063Z digest=sha256:6b2bef6ada38fdeebb61a857c5aa4936a4340242877dbf7ed13820db985ca6ac

Observation 7f87b3e5-0f62-484c-a2bf-6ee9e158b49b · inbound

Representing Speech Through Autoregressive Prediction of Cochlear Tokens cites this paper.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.586186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.586186Z digest=sha256:84cdd961182846e521683b29569ed1f6c1e8d781c59c0aa839106a2ed52be96b

Observation e7a870c0-8ae9-4cb2-b7d2-6361397113fe · inbound

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis cites this paper.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.813060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.813060Z digest=sha256:643a830f74ff0156f2d540c7a3416baa733e081ea0d4cab34a93f51fb14fc13c

Observation 92059eaf-e1d3-4fd6-a32d-58e5de8d0d8c · inbound

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation cites this paper.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.073906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.073906Z digest=sha256:980edbf2279dd8c99f9781414c0fc1d215efd1f6d7bf9c1fb3e4f1a9f3c63476

Observation caa4fed2-b6f6-4b3a-bfee-25ed8bfb1ed6 · inbound

FreeTalk:A plug-and-play and black-box defense against speech synthesis attacks cites this paper.

FreeTalk:A plug-and-play and black-box defense against speech synthesis attacks Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T13:35:04.280612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:35:04.280612Z digest=sha256:fd731b31fe1572338851fc51df50e6789ab547d7384d2375f56d15ec99f4e6be

Observation 801c9947-2627-41c8-acd0-93c70e48918f · inbound

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech cites this paper.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.185099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.185099Z digest=sha256:d9c08f522c276d122ed350f0bfcab88c53cd97ddda9fc21f40f12023cf579029

Observation 1d9d3ebc-8870-446b-8f2e-5308dc2057af · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-21T22:24:23.489484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:6937074b5e9d90396ab8435cf81b82a0ccf591aecf7623f401577198b8fa5e19

Observation 896d6676-1b94-453e-be33-5b385a527b2b · inbound

WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation cites this paper.

WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:01.815189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:36:01.815189Z digest=sha256:359e95bf8f0b68fcf035778d227fb639ab9c202b794a4c8264c4b15f601d343d

Observation 451c10cf-ad20-40b1-81ef-1a87c8c042a6 · inbound

Cloning a Conversational Voice AI Agent from Call\,Recording Datasets for Telesales cites this paper.

Cloning a Conversational Voice AI Agent from Call\,Recording Datasets for Telesales Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T05:51:44.533336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:51:44.533336Z digest=sha256:f1af06dc91bc9975fdeb225c67c2930c78dfe208876c5f071db6e9903c956e0d

Observation 4d66cf36-f702-433f-8c07-6879722d55a0 · inbound

Effectively obtaining acoustic, visual and textual data from videos cites this paper.

Effectively obtaining acoustic, visual and textual data from videos Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:37.053855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:37.053855Z digest=sha256:8dd64c1914e518adc572d7a54937c12c5362ab7f424074aa38833909a73f827c

Observation f489a456-69c3-44f8-9163-9a4a535cf71a · inbound

DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners cites this paper.

DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:53.944636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:53.944636Z digest=sha256:0ae8a1f5f4acf52e90acc16ebaf9cd102845f995b8909f2a08ff0423c200568f

Observation 96f8f981-3f28-4c28-b8f2-5f1b8e04085a · inbound

Testing chatbots on the creation of encoders for audio conditioned image generation cites this paper.

Testing chatbots on the creation of encoders for audio conditioned image generation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-04T21:25:27.309365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:25:27.309365Z digest=sha256:9973015ea62cb97ff8e7e07e21aa60e7176b6e3253001c8e81d01b37988b9ad6

Observation bb4c225b-9529-45bc-b444-56a161c55a5a · inbound

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration cites this paper.

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T19:11:59.593328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:11:59.593328Z digest=sha256:b6907a47d40f70dd79ea2c1cf340dfb49c94c3765bab5cf4bc663e04a387594d

Observation eb2213ca-41f2-4def-9e72-380d921a6b27 · inbound

Length-Aware Rotary Position Embedding for Text-Speech Alignment cites this paper.

Length-Aware Rotary Position Embedding for Text-Speech Alignment Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T17:10:03.685597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:10:03.685597Z digest=sha256:325d6e4907b3094abbb72ee3f9c6ebbec4382fde11924bbe3d542e177cfd6780

Observation 0e51b81a-7193-40de-9309-2975f587cb7b · inbound

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents cites this paper.

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:46:37.601457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T16:44:32.104114Z digest=sha256:1d44ccdfdabe65dec94e4f5c7872879acf938f610c1b82816c1e00d1622906cc

Observation 2de345b8-464a-42fa-ad68-22a579d80986 · inbound

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents cites this paper.

CodecSep: Prompt-Driven Universal Sound Separation on Neural Audio Codec Latents Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T16:48:34.261704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T16:48:34.261704Z digest=sha256:61eb9203ac55da504337953a3f28c89a1acd43828991c5ae8b41e6d58013076e

Observation 2b3821e8-c4cb-4589-b362-5bfc4485c04c · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:36.075318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:36.075318Z digest=sha256:84aa63b04aab8fdf2dd656dfbf4a41608c38cee1166a2e73e87ddf9f417b759c

Observation a7e93bc3-2ecb-4ced-b1d1-41c14daf4228 · inbound

Speculative Coupled Decoding for Training-Free Lossless Acceleration of Autoregressive Visual Generation cites this paper.

Speculative Coupled Decoding for Training-Free Lossless Acceleration of Autoregressive Visual Generation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:20:48.597426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T03:20:35.021120Z digest=sha256:f6c67fdd59cac03187ea5456046972a5d52cb3158dbe1c163e3980529cd6a130

Observation 9289ac55-4351-41a5-85b6-75094684eb58 · inbound

Two-Dimensional Quantization for Geometry-Aware Audio Coding cites this paper.

Two-Dimensional Quantization for Geometry-Aware Audio Coding Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-21T18:20:29.298575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T18:16:51.486807Z digest=sha256:972e5afe322d3e42f8df3c2ef9f2918b5b12236b14d410519bb8bc6413b7cc56

Observation 401bb360-ae68-4c51-844e-693ef14f636f · inbound

Aliasing-Free Neural Audio Synthesis cites this paper.

Aliasing-Free Neural Audio Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:38:24.822590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T20:34:50.539351Z digest=sha256:ba21e0c7ef0f7187562c896f1d9a06b55d53054cc8a879cafa5cc4d184bf740f

Observation c0dad54a-e6e0-4ac5-b440-8a7c80231c7f · inbound

Aliasing-Free Neural Audio Synthesis cites this paper.

Aliasing-Free Neural Audio Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T14:33:09.875302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:33:09.875302Z digest=sha256:fc3f286a01f33e005e06eef565e71fb441b7456dbe870bd54650bc81aacb0680

Observation 407faa63-45d7-4036-a722-81c0e57e3ba1 · inbound

CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models cites this paper.

CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T11:46:12.223223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:46:12.223223Z digest=sha256:397e1b27b7e6ff93ae3cbe3ee673c5c82b1223d396cf4f935c64b07c22fd74f3

Observation 07662bee-7395-480a-923e-d7af4d78f88c · inbound

On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation cites this paper.

On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 4

Resolution
malformed identifier
no resolver link, observed 2026-08-03T11:31:44.746186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:31:44.746186Z digest=sha256:07844a9acb82efef6fca060dc8d905427aa5a62ce2b2ed295e93721260d11dc0

Observation 237646c5-2a90-4e69-977e-307008aef4ea · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:24:56.149288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:6675df6fc902426989164d5fd6d0d15fe75cee16713792f3630e92ed324d4b6f

Observation e58362e4-73da-4412-a91f-d82f8a5896f2 · inbound

SKETCH: Semantic Key-Point Conditioning for Long-Horizon Vessel Trajectory Prediction cites this paper.

SKETCH: Semantic Key-Point Conditioning for Long-Horizon Vessel Trajectory Prediction Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T08:05:55.430857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:05:55.430857Z digest=sha256:914afd2d5c9bb4144568874a7be3e6f25adc3d0b831b470ab28c9395133f3a7d

Observation 0e3f6799-ecbc-415d-ab95-7fdd18c8b98a · inbound

AUHead: Realistic Emotional Talking Head Generation via Action Units Control cites this paper.

AUHead: Realistic Emotional Talking Head Generation via Action Units Control Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:50:40.264150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:49:15.734418Z digest=sha256:5bd03cd7047450bdb76538471c74f7cd354d82a3c88969fa6274e4d35d9f8afc

Observation cc76f225-cba5-4588-b0a1-fb50438b7dc1 · inbound

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model cites this paper.

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T00:11:16.989412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:11:16.989412Z digest=sha256:288e8b600ea571ad0f8509580c27e3ac2ef6130343308556a9bfd8c956e5584f

Observation 0cd1e6aa-b77f-48ac-8e1a-832191215d75 · inbound

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training cites this paper.

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T18:03:40.027879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:03:40.027879Z digest=sha256:1b14b5d526d5de763ab1a16da0eb0492498a0ca477590ab19e647305e5106ce8

Observation 3808e4c4-f71f-4a9c-8fba-5651c4848a1c · inbound

Borderless Long Speech Synthesis cites this paper.

Borderless Long Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:39:50.180256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T07:39:45.386208Z digest=sha256:de7a94d253c88482048205d0d9d3d8619d5a551b5ad63888fc0aa9c7e8e2948e

Observation 59891d00-1594-4f9d-ba16-054ea3478155 · inbound

Voxtral TTS cites this paper.

Voxtral TTS Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T00:39:35.900889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:e4a1bda7dae3798f551e9c25bd413bc4f496ef5a4bf779ed42a56844cea24a27

Observation 04b87d48-ae9f-4f0b-b171-95f4075104c7 · inbound

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck cites this paper.

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:31:07.171839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:51:38.030059Z digest=sha256:35645cb7c0387eaf5e7133b53ef9720ee6083bcb829a6782f8fdc112e3594a94

Observation 8a133291-d5de-4349-9bcc-0ea6f55e40ce · inbound

WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models cites this paper.

WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:29:55.920449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T10:29:51.313835Z digest=sha256:aa97d485275ea93d865423241a133053889980faecd2ce0d69ea1d80be7183a4

Observation 727feb1c-1375-4125-b8b5-28409fa22d77 · inbound

HAFM: Hierarchical Autoregressive Foundation Model for Music Accompaniment Generation cites this paper.

HAFM: Hierarchical Autoregressive Foundation Model for Music Accompaniment Generation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:31:07.171839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:27:56.766226Z digest=sha256:37fdd7691923ea22dbb9d2fa000652b7f98294b2b6e18577a6a02797336dd584

Observation 51985833-c37d-4cbb-bb7b-9174ec1da7ca · inbound

ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models cites this paper.

ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:31:07.171839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:05:45.214298Z digest=sha256:10bbf0b62615870c3be95a82ee9390d17f9ca51ea494c1b1d47aad7d4e9bac0a