Pith. sign in

Paper Citation Record · LEDGER

VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 52 inbound Pith citation observations for arXiv:2406.05370.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.05370 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 52 of 52 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:58:40.546152Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.411558Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 251b1e65-964d-4550-9ac4-6742880289fb · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.535511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:8297cd724ebb4808374f833008fc2cba53841350f52db71f8c7493585ef0e2b9

Observation 5f83f5ca-41ac-4880-8515-d288a2c21a99 · inbound

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models cites this paper.

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:19:09.545037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T06:19:09.507440Z digest=sha256:d2c441dc18c76ede0f0e2ec3feb486c40c28005d9331269487c1ce40e6829a27

Observation 91e21a39-50cf-492f-a160-8d31e37592d9 · inbound

Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate cites this paper.

Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:40.546152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:40.546152Z digest=sha256:d3648e8b740d3ec7f6ff46e886f7ac7ba17d046d6cb3241463c7cfc05c865171

Observation 59084775-ccb8-4363-aad0-0ed7a15b0c55 · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:25.492269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:9201815882f3d6010191b4916bfc20bc2d39650d940ed17605e63325c55686e3

Observation f6446f42-6a4d-4da2-8133-f9a06b7a46a7 · inbound

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling cites this paper.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.602314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.602314Z digest=sha256:d6dbc2bee80db03f0e46ffaf12cd162a3de9e8f9522705a5b3ee4fd579580ac8

Observation 78f0abd3-8987-4043-b332-5830629f06d2 · inbound

Voice Adaptation for Swiss German cites this paper.

Voice Adaptation for Swiss German VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:05.343214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:05.343214Z digest=sha256:0cde56e042dc88fa2a5a4ba5632f750105729e118e252534c54895cc548c5346

Observation 8bf7acdd-960c-4c94-9766-173297f31e72 · inbound

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation cites this paper.

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:25.458653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:25.458653Z digest=sha256:eb891bbfedef7039053925eaf59425f920c0afaca61a64c092b51739b919f137

Observation 170af19d-7e84-4a5d-bd61-fb515caa88a5 · inbound

DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model cites this paper.

DiffDSR: Dysarthric Speech Reconstruction Using Latent Diffusion Model VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:56.370877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:56.370877Z digest=sha256:e51858064d898896e894187b44ddbf16470c574f2296198d7fa38a1681ce18e9

Observation a3105a6d-23ac-44bc-adf1-d7dbaa5fd501 · inbound

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction cites this paper.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:33.743754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:33.743754Z digest=sha256:a92ab98988fa8501758c805e911019bc21163da011e8b33ba8536bc48f867b47

Observation 65de2375-289d-467f-a1fe-5fcfac7c0ea8 · inbound

Speaking images. A novel framework for the automated self-description of artworks cites this paper.

Speaking images. A novel framework for the automated self-description of artworks VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:52.088672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:17:52.088672Z digest=sha256:1bcc0d1b863f3a1634e4808f1d78835705ce481c06794af904a8720a775e77e6

Observation c04f6727-a3cd-412a-8844-60e880dc20d2 · inbound

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model cites this paper.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:11.529192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:11.529192Z digest=sha256:d97d3a70bcb5067e1f7a5f51b3eb7d7e2d0f3b81b23167cc57fdf02479c00e9a

Observation bd381a56-5b71-410b-ad15-d4d22877f2c0 · inbound

Dataset of News Articles with Provenance Metadata for Media Relevance Assessment cites this paper.

Dataset of News Articles with Provenance Metadata for Media Relevance Assessment VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:22.311000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:22.311000Z digest=sha256:55b3203afce740f00d1d4ef697b55b044df90d899c63d30024ea61a742780a20

Observation 564ecf6b-fd82-4b41-b60e-b9bf15fb71bc · inbound

A Variational Framework for Improving Naturalness in Generative Spoken Language Models cites this paper.

A Variational Framework for Improving Naturalness in Generative Spoken Language Models VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.715062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:44.715062Z digest=sha256:37ce4734a20dab8cbad5ad41f1ccf8a5b94dadd8983d60a36c6592035003de67

Observation 9a995078-6c79-4198-9ada-73e1eedc2b32 · inbound

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy cites this paper.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:03.126226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:03.126226Z digest=sha256:21c1d65bbb5d76e57005e51793b0d20c50de6c7d6ac33dd3d68510689f0611c2

Observation c8be1b90-e219-4d50-826e-65bebd9db008 · inbound

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis cites this paper.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.513937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.513937Z digest=sha256:f2498775f1005addd714e11073560e4050a741b4b2e5eb9a7b8a3792526b95f4

Observation 0c650765-46aa-4e7f-868f-ebc596c3d331 · inbound

Next Tokens Denoising for Speech Synthesis cites this paper.

Next Tokens Denoising for Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:25.548417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:25.548417Z digest=sha256:eb2e88268516ebf83eb0908963fa483312d67e24f0a13a78b0d0ad7b095368cc

Observation 53ce4747-da48-485a-b0f5-1a133c8439b7 · inbound

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis cites this paper.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.534536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.534536Z digest=sha256:d171137ef7cede35166f0b2745777a7d96b77df4bb3208d26c63447d71e6927a

Observation 986fe835-d4af-4a6d-886d-bc6c77efe5ff · inbound

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction cites this paper.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:47.001462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:47.001462Z digest=sha256:2dae5d99686533663882f6d91bf8592766f3c6990f9976bba747571809982884

Observation 4dbb91b5-f808-4ded-a2db-5a1266f0324c · inbound

DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching cites this paper.

DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T18:51:20.997584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:51:20.997584Z digest=sha256:2aa8f5b2f66ba21017ae4ffb43089c0198de50f601d8c85ab9bf583df19a13b7

Observation 0b2b2830-f9b6-4d17-90a5-b4d3fc598206 · inbound

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration cites this paper.

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T19:11:59.490276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:11:59.490276Z digest=sha256:1e621e81b0d710afa12bb086d17eea6f347364a533969b1fee0b3fd174b258d0

Observation eaedd33e-6e22-4792-acae-86dc149fecf8 · inbound

CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance cites this paper.

CoMelSinger: Discrete Token-Based Zero-Shot Singing Synthesis With Structured Melody Control and Guidance VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:36:28.820222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T14:34:15.220263Z digest=sha256:4247eee8697f87435eefd095eb433e690e360b1bd7b471e94c2dd4509e681dd9

Observation 0a66343d-5dab-48a9-871d-109ba2cbdaee · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:30.396377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:30.396377Z digest=sha256:0e5399cae84d32d2b5253cecd773150f429024ec0697bebd31d3f95df2c921e0

Observation 35c49ad4-9c77-458f-b191-99ac9d25e756 · inbound

Position: Towards Responsible Evaluation for Text-to-Speech cites this paper.

Position: Towards Responsible Evaluation for Text-to-Speech VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T11:07:39.647731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:07:39.647731Z digest=sha256:db696fa8965ac4883eca1befcc30ff8d1dc5568db55b35a3c7af902efdb335eb

Observation 21fdd2b4-abf1-44b8-a7b8-baf9025ebf8e · inbound

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability cites this paper.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:31.871558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:31.871558Z digest=sha256:1185cf4a1b848b2e9bff1fb7500cef716e376c31a1e1c4e69122aa08e1d7a56b

Observation a0b2c8d6-4306-4cb9-89aa-27d3603f54a1 · inbound

Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection cites this paper.

Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:10:06.050969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T15:07:27.418534Z digest=sha256:c59db7f6685349925eed1aa07b6ea6e0cf82961ae3a5c6d1fd322d5f09143da7

Observation 0e7816cc-9440-42fb-8945-f208a6c3d830 · inbound

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation cites this paper.

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:15:59.846355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:15:28.918204Z digest=sha256:e0416dab0c3b527af2e061a9b7675aff332387a90822b64da75401945c19584a

Observation 6e4a4aa1-4677-401c-a4c4-1a60ac95ac1f · inbound

ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models cites this paper.

ASPIRin: Action Space Projection for Interactivity-Optimized Reinforcement Learning in Full-Duplex Speech Language Models VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:35:58.821060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:05:45.214298Z digest=sha256:715eda49ea516b267b3b0a235e430153835a8c4a362b44edfb7dce29f04a789a

Observation 7f2bb8ca-5d3d-4fda-a860-ad6d9be317d6 · inbound

ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing cites this paper.

ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:21:02.392402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:03:15.572657Z digest=sha256:73f8db09cf8add58d905272f5046afc6c696339f0977dc3f49a5139994db506c

Observation 3bb668d8-ef79-4305-acb6-567cefcafa54 · inbound

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing cites this paper.

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:25.229581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T07:59:34.622096Z digest=sha256:5e65aeba0c7e4a215d3fbff9353e70cf457fcfa541a15a81f4ad2e3bc1a22b46

Observation 85f745ae-e1e7-4840-ac6d-8a4fa94e34ae · inbound

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing cites this paper.

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T05:29:37.288120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:29:37.288120Z digest=sha256:6d8291513f2933ca18d129a083a20e8e8be2cda4d9faecd5512c03f04ca1c793

Observation 45e50250-0aaa-4e0c-90dd-762047a61968 · inbound

SPG-Codec: Exploring the Role and Boundaries of Semantic Priors in Ultra-Low-Bitrate Neural Speech Coding cites this paper.

SPG-Codec: Exploring the Role and Boundaries of Semantic Priors in Ultra-Low-Bitrate Neural Speech Coding VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:11:25.373223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T12:44:08.364835Z digest=sha256:afc0c9c226e445595ef0be3262be924a28146861981d5a80eda58cf4246dec56

Observation 195876d7-3f1f-44ab-8cad-3a89abb7c570 · inbound

Ultra-Low-Bitrate Mel-Spectrogram-based Neural Speech Coding with Flow-Matching-based Refinement and Vocoding-driven Reconstruction cites this paper.

Ultra-Low-Bitrate Mel-Spectrogram-based Neural Speech Coding with Flow-Matching-based Refinement and Vocoding-driven Reconstruction VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:53:55.892033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T19:45:01.705167Z digest=sha256:5160dfec9c1a56c77e0333f7b1397382ecec40a68e71a72de9a8da9d7f20996d

Observation ba4299f2-3240-4434-810f-20192360e422 · inbound

Eroding Trust in Real Speech: A Large-Scale Study of Human Audio Deepfake Perception cites this paper.

Eroding Trust in Real Speech: A Large-Scale Study of Human Audio Deepfake Perception VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:44:48.374423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T15:40:36.466069Z digest=sha256:4adfdee943e11e3a0371e7701c002797391e6877510d57abf3e05ac30e47f78e

Observation 77c95389-f79c-4cd9-95ca-df1846225053 · inbound

Can We Hear from Events? Generating Speech from Event Camera cites this paper.

Can We Hear from Events? Generating Speech from Event Camera VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:43:30.648577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T14:40:29.819278Z digest=sha256:39ec183db1cab2dcc55c50af56d825fcda49f5723e263577adb92ac0a4c89c82

Observation e3030e54-10c4-45dd-9755-9e24c05f2cd1 · inbound

MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables cites this paper.

MELD: Mel-Spectrogram-Based Speech Language Modeling with Discrete Latent Variables VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T05:43:08.711740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T05:34:55.822901Z digest=sha256:6a83c7f07599810496470f7b499c32e9c9c92e2530ff9ba47dd77d4baa5e115d

Observation 5b24562f-9744-4546-9833-a25226c5359b · inbound

UniVocal: Unified Speech-Singing Code-Switching Synthesis cites this paper.

UniVocal: Unified Speech-Singing Code-Switching Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:46:24.605010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T13:17:13.510587Z digest=sha256:427b6155abf2c218f3ec75e42e0918639e43d9abe821dad4589c1377b152d9ba

Observation 16dedabf-4f7f-4a58-8162-5d6a50496af0 · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.798400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:2af45e065c5e7f258fc44fce4d361f2ffe47831a574b5f95af7ba027e6fed1e1

Observation 9350765f-4c89-45dc-affd-64f9f4f2041b · inbound

UniVoice: A Unified Model for Speech and Singing Voice Generation cites this paper.

UniVoice: A Unified Model for Speech and Singing Voice Generation VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:17:08.326755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T23:56:14.198308Z digest=sha256:4baf5eeefd388d1086b3255ecb36792d5bef80d2327f7cbae762f41075779817

Observation 298ce71c-92cb-475a-911c-c4223f441028 · inbound

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech cites this paper.

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:07:35.864020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T15:27:01.142747Z digest=sha256:6548f1a6d7021302ff7ff963b40fc748812e45ef71d9ea4ca6669e959aad0a6b

Observation 054d7587-0a6b-4ae9-971a-4feb9c37c6a1 · inbound

MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data cites this paper.

MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:03.865644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T22:29:05.130421Z digest=sha256:941363ecec930447c61da2e7ef8ea6523e04745281fef6ef342b05a74a80e02d

Observation ca50811b-9982-4aba-a2e4-56b9bda8a131 · inbound

Low-Burden Data Augmentation for Dysarthric ASR via Zero-Shot Voice Cloning cites this paper.

Low-Burden Data Augmentation for Dysarthric ASR via Zero-Shot Voice Cloning VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:19:35.346427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T16:05:40.102418Z digest=sha256:200b1115eeead45b18cb65146ae92cf6359cdd9244771b957b47b24ba7bfc017

Observation a8b77bdb-d009-464a-840b-7cd5be7d4fbb · inbound

PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors cites this paper.

PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:39:40.419205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T15:43:35.759550Z digest=sha256:cf1b5523ea91da6ba787354ffcbe9f06e7c58080c4ddd42180297efbf9330099

Observation ccf1f974-c7e1-4f05-bc22-628c2e7525fc · inbound

Transcript-Free Flow-Matching Text-to-Speech via Speech Feature Conditioning cites this paper.

Transcript-Free Flow-Matching Text-to-Speech via Speech Feature Conditioning VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:39:40.774406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T15:38:13.819676Z digest=sha256:f46e2c0dd1030527089ed93edfe04299686b2488bc72e89c7d9a7c9bbfb80c59

Observation b86b6ea6-9515-4274-bc64-37edf17ea2f8 · inbound

CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations cites this paper.

CrossAccent-TTS: Cross-Lingual Accent-Intensity Controllable Text-to-Speech via Disentangled Speaker and Accent Representations VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:07.163095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T20:21:47.681999Z digest=sha256:8e5826d99daebc76ad0877c6e012b54286f527f0aa2e891d8911029be54aff9a

Observation 1c9e99ca-6521-446e-bb85-08545df4436f · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 116

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.862333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:266913bed264808a029bd9c6ed7a757c58e538c3b8a803c3699c04fe371090c1

Observation 48ea64e8-2bb3-48f4-8f5a-65b0cecc7052 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 236

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.412990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:a86ed945a53f39a269167a1db1771ad99b33e44cb5cdc9c53391a775f8329835

Observation a256841e-f3a6-48a0-8633-cfba7b2c8117 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 236

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:a1c0dbd63a1fc2ae3fd091a8e81fb1399563feb753f9d7c121b635c65a074864

Observation 17ef13f6-c1e3-4fa9-bdf7-dc79538eb0b0 · inbound

SSTMark: Robust Training-Free Semantic-Level Speech Watermarking cites this paper.

SSTMark: Robust Training-Free Semantic-Level Speech Watermarking VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T17:37:38.794755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:37:38.794755Z digest=sha256:db6831fffbf6bc9869cdda1f7474b4d5f838829999eb3bc82d7a0f32a51f6b1f

Observation 41cc9bbd-15a6-4a9d-b01c-e1eafaa25407 · inbound

Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English cites this paper.

Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T03:52:17.151139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:52:17.151139Z digest=sha256:206c956bb93cc0538937fc1b2387a48bf23c13bdc5988a0703f3ecb8044838ae

Observation ce2fba8b-b260-40ef-a172-decdda9b2d4c · inbound

Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis cites this paper.

Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T15:08:47.779386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T15:08:47.779386Z digest=sha256:de49835259fbc729d3a37464e7e16a69ab0da3a09d4f2b4a3447ca7a50dff580

Observation 7878a799-da21-4452-a938-411c8de1e7f4 · inbound

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens cites this paper.

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T08:35:47.644694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:35:47.644694Z digest=sha256:9dc8437629a0d985d9bc67b78ecd6d9d9aabc02f3a24cceeee7f6ae8de5e3175

Observation 3ee3d8f0-ae27-4ffb-a62f-2f41c2b2fb4c · inbound

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis cites this paper.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:35.142024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:35.142024Z digest=sha256:29b22f4d75f2051a0a9ad2eca9ca3a5887307e38e8783737a17e3fccb889d581