Pith. sign in

Paper Citation Record · LEDGER

Next Tokens Denoising for Speech Synthesis

As of 17 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2507.22746.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22746 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:22:27.950994Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact5
  • verified fuzzy1
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6229081c-1426-49c8-90a8-da36704ba759 · outbound

This paper cites Better speech synthesis through scaling.

Next Tokens Denoising for Speech Synthesis Better speech synthesis through scaling

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:25.341706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:25.341706Z digest=sha256:1a126959827c7e9bde6de2f208c27de5d4a74cb0c80b968bcba8e3bf15811b1c

Observation 3b306d4d-dbb8-43b9-b58d-7acdf6d27af1 · outbound

This paper cites High Fidelity Neural Audio Compression.

Next Tokens Denoising for Speech Synthesis High Fidelity Neural Audio Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:25.630747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:25.630747Z digest=sha256:d3af81e276ce96818333c04f7da853dbb6f574d7178e20d12ba6e52a8367d0d2

Observation 4c112359-d552-43bb-9d01-693ec47a680d · outbound

This paper cites Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback.

Next Tokens Denoising for Speech Synthesis Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:25.922271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:25.922271Z digest=sha256:33a8cae0c009d27ce81b09add35753db85c69c7892a6ce9f69963b37ca7d0d0b

Observation 3265b4f7-b4a9-463f-9301-ac12a400f9da · outbound

This paper cites Mean Flows for One-step Generative Modeling.

Next Tokens Denoising for Speech Synthesis Mean Flows for One-step Generative Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:25.951018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:25.951018Z digest=sha256:d1a165439482ea5818249c6532a34714e5df33059fac8b1864bd8d6b7b8fc3b5

Observation ce0731a6-61e7-4eb0-bc97-2580a1e39188 · outbound

This paper cites Ditar: Diffusion transformer autoregressive modeling for speech generation.

Next Tokens Denoising for Speech Synthesis Ditar: Diffusion transformer autoregressive modeling for speech generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:26.029534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:26.029534Z digest=sha256:25db8d85b11612551eefcc94a533c2d380c9c81919d46977b0b970e3692f7643

Observation 42b581fd-1e21-41bf-b2d1-564f140f6444 · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

Next Tokens Denoising for Speech Synthesis NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:26.095836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:26.095836Z digest=sha256:720cca6167735e06f3b99d2fdb640d051702a5aaf4b2b429a67cd0a629782e7a

Observation a5fa5a1c-84dd-47af-8e12-f28339bbaec8 · outbound

This paper cites Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like.

Next Tokens Denoising for Speech Synthesis Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:26.208323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:26.208323Z digest=sha256:d88c4f29d07626191b72b3deb70dfbb92fdf804869dbef2b233a438203da851b

Observation aa2df25a-9a0d-4912-8631-37d282ff4376 · outbound

This paper cites istftnet: Fast and lightweight mel-spectrogram vocoder incorporating inverse short-time fourier transform.

Next Tokens Denoising for Speech Synthesis istftnet: Fast and lightweight mel-spectrogram vocoder incorporating inverse short-time fourier transform

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:22:29.105049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:22:26.324190Z digest=sha256:4890c4ec2a1481f84365f5eaf67b4c2f1933b507f331860033d3d7daf70a25e6

Observation 16933749-f487-4e28-8067-c097a54726a6 · outbound

This paper cites Understanding DDPM Latent Codes Through Optimal Transport.

Next Tokens Denoising for Speech Synthesis Understanding DDPM Latent Codes Through Optimal Transport

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:26.402690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:26.402690Z digest=sha256:ffb599fb23e78240c3ee082cc210b04c8cdb6c0b2b9eaf77e7e38a05ca44a020

Observation 3763f71c-b841-4b96-8274-d7fe01dee4d1 · outbound

This paper cites PromptTTS 2: Describing and Generating Voices with Text Prompt.

Next Tokens Denoising for Speech Synthesis PromptTTS 2: Describing and Generating Voices with Text Prompt

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:26.482050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:26.482050Z digest=sha256:751fba14ce7cc38aafaedf1a6897caeaad7e41dd4dcbbaf884120cafc6dc1d2c

Observation 2d8b244c-560b-408d-a460-e5e525c30581 · outbound

This paper cites Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation.

Next Tokens Denoising for Speech Synthesis Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:22:28.855073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:22:26.545205Z digest=sha256:1106162181c8176570fc1b7bd8c6566b9c6899231c2959ec4e577e45666cf0c4

Observation 1877fbd7-3822-4a5d-95c0-d4b452f46c9f · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Next Tokens Denoising for Speech Synthesis Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:26.665411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:26.665411Z digest=sha256:69deb21469bb247660bfd57d9e81645ad6f849324ddd5ce15c582264749e19b5

Observation 5c6132fe-1c44-4db0-a7a9-dc422b4ccf88 · outbound

This paper cites DelightfulTTS 2: End-to-End Speech Synthesis with Adversarial Vector-Quantized Auto-Encoders.

Next Tokens Denoising for Speech Synthesis DelightfulTTS 2: End-to-End Speech Synthesis with Adversarial Vector-Quantized Auto-Encoders

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:22:28.570689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:22:26.737905Z digest=sha256:85dac6ff21816094aab11cbcb9e8d6732864619d9a19c586c0472c7745884d41

Observation d3549354-974d-4d53-b100-5142f17f4f53 · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

Next Tokens Denoising for Speech Synthesis Autoregressive Speech Synthesis without Vector Quantization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:26.796662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:26.796662Z digest=sha256:993a6fa7290f887f895c1fa13be47232c162bf1d9281548173aa6b8ebff32797

Observation d4b868e7-0546-4ed3-be84-a08794a3fc41 · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

Next Tokens Denoising for Speech Synthesis Finite Scalar Quantization: VQ-VAE Made Simple

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:26.883167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:26.883167Z digest=sha256:7fe9d587ff38fe7d46d7990bc430e7dba0144916399099174f1d67de3eb5f031

Observation f01925d6-4100-4867-b4ef-5eb009bfc73c · outbound

This paper cites Scaling Transformers for Low-Bitrate High-Quality Speech Coding.

Next Tokens Denoising for Speech Synthesis Scaling Transformers for Low-Bitrate High-Quality Speech Coding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:26.987973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:26.987973Z digest=sha256:e659acec6fe7ca476eab83a1f730baad0aeb2879797e4e9bb7d94790ac993365

Observation d123e31a-3705-44d4-864b-225d16ba3970 · outbound

This paper cites Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion.

Next Tokens Denoising for Speech Synthesis Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:27.053403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:27.053403Z digest=sha256:862413a53c12a8caad5af7ea1a7cb3e3de86fa02146501f28d79e5e073a15205

Observation faba2588-9e08-487e-b264-81415b3455e9 · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

Next Tokens Denoising for Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:27.121148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:27.121148Z digest=sha256:08e3ded400f0603dbb067ee7bf13bbacd44ec91a30b01463f27f407cb5fcac42

Observation 4706dce0-0b29-472a-80c5-38080d13a8c1 · outbound

This paper cites Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling.

Next Tokens Denoising for Speech Synthesis Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:22:28.411634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:22:27.182862Z digest=sha256:217a5d696ed99760e202b498c25d766940c45742355bdb68e0b40881cc57f653

Observation d503ec9f-d481-4f10-aab6-d5a8b36a0787 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Next Tokens Denoising for Speech Synthesis Gemini: A Family of Highly Capable Multimodal Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:27.280963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:27.280963Z digest=sha256:e244c7756e4137e13c0f84a0a8312f175801140440fb2b8e9d477f93fa9ff151

Observation 7001ad77-5a82-47ef-a1b7-01fb19ad16d9 · outbound

This paper cites Improving and generalizing flow-based generative models with minibatch optimal transport.

Next Tokens Denoising for Speech Synthesis Improving and generalizing flow-based generative models with minibatch optimal transport

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:27.357575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:27.357575Z digest=sha256:d6077484974695d1975fda86ca751b74ffe3635124e162e621ec5c52be4f166b

Observation efbba303-682e-4bdb-9678-fe4e6a456f05 · outbound

This paper cites FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching.

Next Tokens Denoising for Speech Synthesis FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:27.493156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:27.493156Z digest=sha256:d50c168dfb6eaf0f3912e9fa308b0377f592d7169c01a51fb6c2b76c2f39f0c6

Observation b2a6f30a-9185-4d39-bc60-d608c1c65b1c · outbound

This paper cites Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis.

Next Tokens Denoising for Speech Synthesis Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:27.528989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:27.528989Z digest=sha256:233b8aa9f77fb41a05d15b89dde3b84f8f659853694c904fba8be10c04dc8cb5

Observation f3d6bb98-fc36-488c-8faf-0489bb1a1767 · outbound

This paper cites Lumos-1: On autoregressive video generation from a unified model perspective.

Next Tokens Denoising for Speech Synthesis Lumos-1: On autoregressive video generation from a unified model perspective

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:27.625768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:27.625768Z digest=sha256:28c76f5c9b194c6d3bca2c01f870231780a49fae67bee7d2e2fc67af19ef9ebe

Observation 57703305-fd3c-4bd1-8837-dd90a5334c5e · outbound

This paper cites Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners.

Next Tokens Denoising for Speech Synthesis Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:27.693175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:27.693175Z digest=sha256:1255628a1d8ee467ae83a2a4459044939bb85bcb0cccebcdc5a578e9924e61d8

Observation 12782ae2-731c-4d75-bc15-cfd92c294d8f · outbound

This paper cites Boosting Diffusion Model for Spectrogram Up-sampling in Text-to-speech: An Empirical Study.

Next Tokens Denoising for Speech Synthesis Boosting Diffusion Model for Spectrogram Up-sampling in Text-to-speech: An Empirical Study

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:22:28.145958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:22:27.789945Z digest=sha256:bbc8ee0fedf8e9189cc10850d69862c5002be21057cb8883e7ea1cb74d27c2ea

Observation 53ce42bb-f56c-4f05-b7e7-2ad9bf1494fd · outbound

This paper cites Mixed-Phoneme BERT: Improving BERT with Mixed Phoneme and Sup-Phoneme Representations for Text to Speech.

Next Tokens Denoising for Speech Synthesis Mixed-Phoneme BERT: Improving BERT with Mixed Phoneme and Sup-Phoneme Representations for Text to Speech

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:27.883606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:27.883606Z digest=sha256:35c78492bab206f2f9c47cf295e6611c79c8cbe7f7213566ebd1a9273b3243d6

Observation d4517a79-8173-4f65-8551-ffcd886f9f98 · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

Next Tokens Denoising for Speech Synthesis Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:27.950994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:27.950994Z digest=sha256:bc097eda649d54ac6b9228316bd83057c15fdf533b121b70b13bc693160bee69

Observation 2e844282-1ca3-4230-b734-949e062f358f · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Next Tokens Denoising for Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:27.428195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:27.428195Z digest=sha256:286943086508f46cc6c568354d38029ada6543a50b316f464a73b267bc873e92

Observation ebfe630b-423b-4645-a172-64b64bff40ab · outbound

This paper cites MoBoAligner: a Neural Alignment Model for Non-autoregressive TTS with Monotonic Boundary Search.

Next Tokens Denoising for Speech Synthesis MoBoAligner: a Neural Alignment Model for Non-autoregressive TTS with Monotonic Boundary Search

Reference 2019

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:22:28.701598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:22:26.621996Z digest=sha256:0149e8d5a63ab85f54daa7e84d0294e8fe48424ed5a2cde7bbbdadf3611a66c8

Observation d8d21cbe-36a0-4a83-ae5c-37f0360f7ef3 · outbound

This paper cites AdaSpeech: Adaptive Text to Speech for Custom Voice.

Next Tokens Denoising for Speech Synthesis AdaSpeech: Adaptive Text to Speech for Custom Voice

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:25.496315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:25.496315Z digest=sha256:1762d22841f9874f1da2fe000c1dd3b568f32ea93938b1a0537c46d79cb085ab

Observation 0c650765-46aa-4e7f-868f-ebc596c3d331 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

Next Tokens Denoising for Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:25.548417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:25.548417Z digest=sha256:a1e2de5c2972539830a1ddaa7e8326336b792df89ae0f78b0599d54220e3e5e6

Observation 572e330a-4104-4d2b-b35e-f211c1787355 · outbound

This paper cites E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS.

Next Tokens Denoising for Speech Synthesis E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:25.706245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:25.706245Z digest=sha256:6c29fae1add11e3b7479dd3b31735a4d1feafc9ad2dcc7437a0230a22625aa23

Observation 9e9ba83c-6772-4838-9f60-dc93aea1ba61 · outbound

This paper cites Language models are few-shot learners.

Next Tokens Denoising for Speech Synthesis Language models are few-shot learners

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:25.418317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:25.418317Z digest=sha256:19697cbc18e7830e49d53b868ecc91621226c50378da0823060539acb2ebe70b

Observation 68f10129-0acc-4fc5-a3d2-d6a4e1b0060c · outbound

This paper cites One Step Diffusion via Shortcut Models.

Next Tokens Denoising for Speech Synthesis One Step Diffusion via Shortcut Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:25.814243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:25.814243Z digest=sha256:6ab92cdad59002c67ec92908d553a2d39b66838f0b13d512597d3d2a7848b5b2

Observation 0b58e0cd-015b-4b9c-a206-ecd55188d95f · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

Next Tokens Denoising for Speech Synthesis VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:25.974745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:25.974745Z digest=sha256:4fdf7c50fd6618b1587313f585e1c0a90f86a16f75a8cd3ddd3f29e494f95940

Pith citing papers

No inbound Pith citation observations are available.