Pith. sign in

Paper Citation Record · LEDGER

Luna-TTS Family Technical Report

As of 22 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 0 inbound Pith citation observations for arXiv:2608.11593.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11593 v1

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:40:25.380445Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

88 of 88 outbound references displayed

  • verified exact8
  • verified fuzzy24
  • unresolved55
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e805ec3a-4354-44c9-8abb-983e9e7aaae1 · outbound

This paper cites SoundStream: An end- to-end neural audio codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507, 2022.

Luna-TTS Family Technical Report SoundStream: An end- to-end neural audio codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507, 2022

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.887133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.887133Z digest=sha256:1bd91c9e1ca40091ae28c274b0995a00952a692715b7ca15ad9ce55507daa0cc

Observation f48e3008-64ad-4001-878c-7fd84a817a87 · outbound

This paper cites High Fidelity Neural Audio Compression.

Luna-TTS Family Technical Report High Fidelity Neural Audio Compression

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.893723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.893723Z digest=sha256:a31aa12a3c0791ea1524db8be86ece2a5c454294769175c4ff5cb238e93b00c6

Observation b9b46316-33a2-415e-9ff5-5d97a84632e4 · outbound

This paper cites High-fidelity audio compression with improved RVQGAN.

Luna-TTS Family Technical Report High-fidelity audio compression with improved RVQGAN

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.899526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.899526Z digest=sha256:16beddf3a13fc9201ee7876425966b0f484e11febc40fe15e170ddb72fe7a127

Observation f33112ad-8a97-4523-85de-78b437ba2705 · outbound

This paper cites Better speech synthesis through scaling.

Luna-TTS Family Technical Report Better speech synthesis through scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.905147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.905147Z digest=sha256:6b17cec4388c7263494a33f2b837df289e6153cfa8049f6fed0362f90e883657

Observation c3011b86-331d-4869-8ea6-355dcd2ef766 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Luna-TTS Family Technical Report Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.911049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.911049Z digest=sha256:ddf4e8e3f4e6d69d64fa4c815f2ebec210aa1e56865e59b8694ba8f33aaae365

Observation ff403716-cafb-4b0e-a78e-9d0bb87ed7db · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Luna-TTS Family Technical Report Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.916854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.916854Z digest=sha256:15b0bc951e89512b974f02d344ccbd12b03b9c70ef44d76fc21d0bf51521a0a2

Observation 389b2010-ca7c-484a-aee2-1dd823fad805 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Luna-TTS Family Technical Report CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.923060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.923060Z digest=sha256:975ab2079d22d6229625d644835113198f700129162c61557ab33d2a2efab299

Observation bb718ee1-7cd7-4fd5-86e5-ee0a231c36f0 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Luna-TTS Family Technical Report CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.928185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.928185Z digest=sha256:edfb61695af42d9d91861688fccc76e649d9e185e7c756790fc63612ffc6a713

Observation 128799c2-0bda-4d2e-8bd3-85b2e53ae1d4 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

Luna-TTS Family Technical Report CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.933442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.933442Z digest=sha256:6a9e6309ca1542cefa6bf01079e515287c958b8529600340a07d5755336ef655

Observation 7dd17738-c6e4-496b-8b23-6700317c5213 · outbound

This paper cites MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder.

Luna-TTS Family Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.938990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.938990Z digest=sha256:2da361e33dd80372be2527a564a860f9bb3adf6dab007cdcdf42030d4072b399

Observation a247d038-f48d-43b7-8843-9a0cabb861a1 · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

Luna-TTS Family Technical Report Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.944475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.944475Z digest=sha256:6682afa4237b03b20721c36f814ae1b9a5f4d18d8e43fe5abfbda1e9c5cb6e19

Observation ba110c96-b42a-40f7-a287-93fc4fb0014a · outbound

This paper cites GLM-TTS technical report.arXiv preprint arXiv:2512.14291, 2025.

Luna-TTS Family Technical Report GLM-TTS technical report.arXiv preprint arXiv:2512.14291, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.949797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.949797Z digest=sha256:ce8574e35d0c5f7633e24c6a0857f66a2a7a4fab8c59b6ebb8937765c6f44544

Observation d7b2c8f1-4399-4437-a100-28f26ec637fc · outbound

This paper cites Qwen3-TTS Technical Report.

Luna-TTS Family Technical Report Qwen3-TTS Technical Report

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.954815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.954815Z digest=sha256:549ea0eabc2fca4aa0aa6b63d4c29b39371d5c114f1fdd90f250d1021241ad68

Observation a15fa500-ebc6-46ef-9891-7ce5705d6651 · outbound

This paper cites Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm.

Luna-TTS Family Technical Report Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:40:26.847896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:24.960358Z digest=sha256:a2aae594cbf9a061614928225064ff3e74c6fe5a1502607389186aa0a933a96e

Observation 87d36c2d-eb45-476c-8252-4cc375bc3c07 · outbound

This paper cites Fish audio S2 technical report.arXiv preprint arXiv:2603.08823, 2026.

Luna-TTS Family Technical Report Fish audio S2 technical report.arXiv preprint arXiv:2603.08823, 2026

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.965356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.965356Z digest=sha256:6c56e06d80689d19127587b62c8fc9708080a7cff144ecce19c7c7fcf9a75b8c

Observation 89ad41b8-bb8e-48ea-bf2d-f69c172c8371 · outbound

This paper cites MOSS-TTS technical report.arXiv preprint arXiv:2603.18090, 2026.

Luna-TTS Family Technical Report MOSS-TTS technical report.arXiv preprint arXiv:2603.18090, 2026

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.970636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.970636Z digest=sha256:e8349e8d12e5eb19f4ec104273ada0a25a362bf0498685f901b7897235bc76f6

Observation 2000ce91-14ae-4b79-9ae6-bb88bd612383 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

Luna-TTS Family Technical Report Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.975736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.975736Z digest=sha256:f4771001c12d1ddc89bcfdd3ba5f936692fb91bd287f3b967614b30b85511698

Observation 2a4a3f19-edd3-4042-8c13-52514f576082 · outbound

This paper cites IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech.

Luna-TTS Family Technical Report IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.981589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.981589Z digest=sha256:57674fd8c2929d47856101ee3e0b92f20cbda9b7cc57de64e54c819870fa430e

Observation f801373a-e6a9-4e29-94d6-3317bd77ead8 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

Luna-TTS Family Technical Report FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.987167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.987167Z digest=sha256:3e93d90bf6691ec7e49a87240759ea38b0bbf7733db3726dbeec642f78060727

Observation 5e314825-09b3-4d1b-a1b0-634f1246503e · outbound

This paper cites MiMo-Audio: Audio language models are few-shot learners.arXiv preprint arXiv:2512.23808, 2025.

Luna-TTS Family Technical Report MiMo-Audio: Audio language models are few-shot learners.arXiv preprint arXiv:2512.23808, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.993372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.993372Z digest=sha256:6fac0f904de6676229f8be79419ee40441defc8b3fac6583a383fa51837f7a56

Observation 2cb6cdce-fade-48dc-9163-1355ecf94a81 · outbound

This paper cites Step-Audio 2 Technical Report.

Luna-TTS Family Technical Report Step-Audio 2 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:24.998635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:24.998635Z digest=sha256:f413df87dbc2c7431a5a27eba4e7686d08e8bcbc06cc781867961504f42a7b14

Observation 667be8b7-5893-4834-8b1d-3296326292ab · outbound

This paper cites Simple and controllable music generation.

Luna-TTS Family Technical Report Simple and controllable music generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.004435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.004435Z digest=sha256:ad12121e6d7cfdd5a80a06498d65d36ea484daa08dffdca3deef2904cf9b1ffd

Observation 8b86a200-3300-4270-be8d-5eb160fd1cf6 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Luna-TTS Family Technical Report Moshi: a speech-text foundation model for real-time dialogue

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.009906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.009906Z digest=sha256:3f577bc8ae1a9456a55a014a731e6ff9d329c1a7fe416e2d4627551c68ab6819

Observation 3d139120-1b51-46bd-9772-eb3b502008af · outbound

This paper cites DiSTAR: Diffusion over a scalable token autoregressive representation for speech generation.arXiv preprint arXiv:2510.12210, 2025.

Luna-TTS Family Technical Report DiSTAR: Diffusion over a scalable token autoregressive representation for speech generation.arXiv preprint arXiv:2510.12210, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.015753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.015753Z digest=sha256:65ddce81a2869ec80f11ed738185c8f2c9eb361fce4676f512e695ebad92a626

Observation 6aad38f3-6432-4f68-a72a-bb4cac0dfa31 · outbound

This paper cites an unresolved cited work.

Luna-TTS Family Technical Report Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.022798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.022798Z digest=sha256:d6e281be3311f7196416c9dea4c6a3513ab1e510c21e4f78a5319be1c8956de0

Observation 87fbe408-4a65-4dac-8110-c350f42be3cc · outbound

This paper cites V oicebox: Text-guided multilingual universal speech generation at scale.

Luna-TTS Family Technical Report V oicebox: Text-guided multilingual universal speech generation at scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.028068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.028068Z digest=sha256:abdae79e273d80f0a8c15877b211d2bffea9c61d384a31272c380965b7a5fdfd

Observation 0db4ed6c-0757-4460-931c-660c015f3bec · outbound

This paper cites E2 TTS: Embarrassingly easy fully non-autoregressive zero-shot TTS.

Luna-TTS Family Technical Report E2 TTS: Embarrassingly easy fully non-autoregressive zero-shot TTS

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.033611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.033611Z digest=sha256:d3f8a16257f95c27067a0dffd783e20800b237a9365587a38c822fa7604e6cf6

Observation cd6fe3c4-7543-4c15-a2d8-0b290ccd39a4 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

Luna-TTS Family Technical Report F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.039576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.039576Z digest=sha256:2354e006b0c31a9876ec6a6ac024609e0585349bca57405c130aec03df938e0a

Observation 0ea104cb-4681-4bef-8298-85341d3e443a · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

Luna-TTS Family Technical Report SoundStorm: Efficient Parallel Audio Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.044878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.044878Z digest=sha256:4b19591ea3dbcd9b8d359543b2998f7731007473961e551cfdfb9f56905d1612

Observation a46bc923-c6f9-4d8f-bd1d-9b285d13d33d · outbound

This paper cites NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models.

Luna-TTS Family Technical Report NaturalSpeech 3: Zero-shot speech synthesis with factorized codec and diffusion models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.051412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.051412Z digest=sha256:425c63fa7f8da5a2eaf7f11b643838d910caeaeb3e460216070af1d52da58eb6

Observation f26c9346-c6c8-4b2e-898f-b88d471880ba · outbound

This paper cites MaskGCT: Zero-shot text-to-speech with masked generative codec transformer.

Luna-TTS Family Technical Report MaskGCT: Zero-shot text-to-speech with masked generative codec transformer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.056917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.056917Z digest=sha256:392007ba2628076b26fb2b46e73674df52479a266b3776b02acfc9075acab0b8

Observation d2f22061-83bf-47ce-b889-d188dcf62b77 · outbound

This paper cites Johnson, Jonathan Ho, Daniel Tarlow, and Rianne van den Berg.

Luna-TTS Family Technical Report Johnson, Jonathan Ho, Daniel Tarlow, and Rianne van den Berg

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.732705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.062469Z digest=sha256:f227a65e9de1acb1bef794a6f9da6325649d11403946dcc05a8ab530a18bee7f

Observation db57c669-3b5a-41b7-abd5-d0cd2eebf058 · outbound

This paper cites Discrete diffusion modeling by estimating the ratios of the data distribution.

Luna-TTS Family Technical Report Discrete diffusion modeling by estimating the ratios of the data distribution

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.067949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.067949Z digest=sha256:031a2c076c45d66dfbc1f8baa5241ce61c9269af20ad4155318f6f384ac3b08e

Observation f8b29698-a3a1-485e-96de-583cf3fe2979 · outbound

This paper cites Chiu, Alexan- der Rush, and V olodymyr Kuleshov.

Luna-TTS Family Technical Report Chiu, Alexan- der Rush, and V olodymyr Kuleshov

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.697054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.074048Z digest=sha256:3ad663b4068ed5051b42e4959389d4da2e35c7bd0536c28b553562e462367b91

Observation f40f2ebe-67dc-442b-86dc-4e313bfd3047 · outbound

This paper cites an unresolved cited work.

Luna-TTS Family Technical Report Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.079322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.079322Z digest=sha256:a76f2b313670361553b6bedca60d5947534bc6145fb6f2e9139f1a93d9bbce79

Observation 2f13b8b3-cca3-43a4-b397-4f250a409bb6 · outbound

This paper cites Large Language Diffusion Models.

Luna-TTS Family Technical Report Large Language Diffusion Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.084430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.084430Z digest=sha256:8bea64f3359efd30a1a9224f2df0da978f4690f0dab419319977daf6d9cfa24f

Observation bcfce8e3-3daf-40e2-89af-aa3f9aff2a39 · outbound

This paper cites Dream 7B: Diffusion Large Language Models.

Luna-TTS Family Technical Report Dream 7B: Diffusion Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.090799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.090799Z digest=sha256:73e8bfaeccd67a8c97cda8176e23fdb69d942a63022ec863925300e7113dfe08

Observation 9202085e-1b06-4928-9a34-f94072035da1 · outbound

This paper cites Mercury: Ultra-Fast Language Models Based on Diffusion.

Luna-TTS Family Technical Report Mercury: Ultra-Fast Language Models Based on Diffusion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.096604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.096604Z digest=sha256:61818bd1434d14ddffc7fcf103c065bcea96b8d874fb26f740a533a947f6dae7

Observation 2fa27e5d-7ef2-41df-a20b-8e8ac16fec95 · outbound

This paper cites Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference.

Luna-TTS Family Technical Report Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.102291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.102291Z digest=sha256:e05347d8554f4f4c15e9738945d8bb55bde9e17ca5fe3539146a6679a7e0f26e

Observation f1b2993f-3e95-4249-abfd-acd88a730c93 · outbound

This paper cites LLaDA2.0: Scaling Up Diffusion Language Models to 100B.

Luna-TTS Family Technical Report LLaDA2.0: Scaling Up Diffusion Language Models to 100B

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.109186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.109186Z digest=sha256:a0d0edfb6bd5af3fb99d07b7287b2d7982640ed9f79296b7e3b9bab9ecfdfd7b

Observation 591840f3-f639-4be5-9501-f462409acf9d · outbound

This paper cites Chiu, Zhihan Yang, Zhixuan Qi, Jiaqi Han, Subham Sekhar Sahoo, and V olodymyr Kuleshov.

Luna-TTS Family Technical Report Chiu, Zhihan Yang, Zhixuan Qi, Jiaqi Han, Subham Sekhar Sahoo, and V olodymyr Kuleshov

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.665634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.114717Z digest=sha256:b49f7f5ab14bf6e0e4855ecbd19882ecebdae9587b89559888d8a82612e62fbe

Observation 478ede88-ccb1-4661-8c2b-af6f3ed851d8 · outbound

This paper cites Scaling diffusion language models via adaptation from autoregressive models.

Luna-TTS Family Technical Report Scaling diffusion language models via adaptation from autoregressive models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.646767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.121603Z digest=sha256:16ae08af88955412bc7094ece44573e99ae55fde9f14104479622086adebd40c

Observation afe55eeb-5f63-4f55-a599-22f2c2063eb6 · outbound

This paper cites Fast-dLLM v2: Efficient block-diffusion LLM.

Luna-TTS Family Technical Report Fast-dLLM v2: Efficient block-diffusion LLM

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.128873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.128873Z digest=sha256:49072ec72c58319fdf7643d2f51392fd4fd89aea43c334c35c26ff073dde73a8

Observation 00a6024b-3379-447b-966e-e8aef4bcf578 · outbound

This paper cites Sequential diffusion language models.arXiv preprint arXiv:2509.24007, 2025.

Luna-TTS Family Technical Report Sequential diffusion language models.arXiv preprint arXiv:2509.24007, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.134673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.134673Z digest=sha256:4fbf69f1ad35de1a70843393d8b5ec456b391c9ea5da31be8aa87b14b8726ae2

Observation 78e4914f-fb63-468c-98b7-979a34a62d7c · outbound

This paper cites StepAudio 2.5 Technical Report.

Luna-TTS Family Technical Report StepAudio 2.5 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.139841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.139841Z digest=sha256:46ce7cfe9ee7f4bb941ddf5910274389b14fcf803e052ba67dc529dd3c688c7f

Observation 57f94976-6046-443a-b9c1-dc5fdd30c87c · outbound

This paper cites LLaDA-TTS: Unifying speech synthesis and zero-shot editing via masked diffusion modeling.arXiv preprint arXiv:2603.26364, 2026.

Luna-TTS Family Technical Report LLaDA-TTS: Unifying speech synthesis and zero-shot editing via masked diffusion modeling.arXiv preprint arXiv:2603.26364, 2026

Reference 46

Resolution
verified exact
raw_fallback, observed 2026-08-16T00:40:26.111805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.145077Z digest=sha256:9c35ddd41f5be8d0c89e7356ae6c3dbcb2be541b666e2a9377423d92a78db4a7

Observation e81d885d-e5f5-41b9-a2d9-4d7891f9e834 · outbound

This paper cites DiffuSpeech: Silent thought, spoken answer via unified speech- text diffusion.arXiv preprint arXiv:2601.22889, 2026.

Luna-TTS Family Technical Report DiffuSpeech: Silent thought, spoken answer via unified speech- text diffusion.arXiv preprint arXiv:2601.22889, 2026

Reference 47

Resolution
verified exact
raw_fallback, observed 2026-08-16T00:40:26.028258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.150006Z digest=sha256:fbde5a5335a91bb7809850819fee424ae69d68a1772433e1d87b17e6fde24c03

Observation 98a5f179-4fed-4619-8621-08280873dec0 · outbound

This paper cites OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models.

Luna-TTS Family Technical Report OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.155751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.155751Z digest=sha256:87d4ed6e746bcd55b17295a05bc1f4d42834fa8560177f2a530fca0e1b3dba9f

Observation 490026c5-ac1c-4f97-aabd-7304ef8c2fe2 · outbound

This paper cites Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS.

Luna-TTS Family Technical Report Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:40:25.918892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.161585Z digest=sha256:df5a9d943aaa10075a933a2f2f87460749151478ea20f3570cc02abe76a9a9ca

Observation 074360d6-8526-4c92-93be-bb04c4490469 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Luna-TTS Family Technical Report DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.166914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.166914Z digest=sha256:588cde4b3780381340e92d7d0a6a7cdae6f25963c708153f82e1af318c19bdde

Observation 52cc8349-ac76-42dd-98c5-46cec00de6c5 · outbound

This paper cites V oxCPM: Tokenizer-free TTS for context-aware speech genera- tion and true-to-life voice cloning.arXiv preprint arXiv:2509.24650, 2025.

Luna-TTS Family Technical Report V oxCPM: Tokenizer-free TTS for context-aware speech genera- tion and true-to-life voice cloning.arXiv preprint arXiv:2509.24650, 2025

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.171837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.171837Z digest=sha256:5a08660df60b79b1604a08bbca0ca3c90db0eaf623a28d665a77058149dc5804

Observation e2a714f9-8a76-4278-8760-ee261ec3b95d · outbound

This paper cites Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model.

Luna-TTS Family Technical Report Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.177667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.177667Z digest=sha256:3d58424ad09f2999c785e285ea69035e4616019c42061023795b01f955b90bbb

Observation 9cebe932-f49a-4378-a694-f4ee21a69bcf · outbound

This paper cites SpeechTokenizer: Unified speech tokenizer for speech large language models.

Luna-TTS Family Technical Report SpeechTokenizer: Unified speech tokenizer for speech large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.628956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.183840Z digest=sha256:ab3cc4f8e4cb024db4ce102f28b2e2f0974b3651a5e5ab0b7584da4085526298

Observation 62de1417-535f-44f5-94d1-18ef2f47568b · outbound

This paper cites WavLM: Large-scale self-supervised pre-training for full stack speech processing.

Luna-TTS Family Technical Report WavLM: Large-scale self-supervised pre-training for full stack speech processing

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.609658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.188805Z digest=sha256:2836472903af4a7b908a39b4ced2a2d47c7e83e2883fea871b8a5bfffc52d97b

Observation 388b7fb9-1796-4bd3-8956-e651984f2b11 · outbound

This paper cites HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis.

Luna-TTS Family Technical Report HiFi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.589174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.194056Z digest=sha256:a56b16d4193ae4f2271b1becb3cfab8940562cc5dd323972aa85021d6e9b40fb

Observation 631bdc40-d89e-4ac7-ace5-31312e8b7fcb · outbound

This paper cites BigVGAN: A universal neural vocoder with large-scale training.

Luna-TTS Family Technical Report BigVGAN: A universal neural vocoder with large-scale training

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.569169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.199085Z digest=sha256:63acb8f18d21473eb4f6f614818305808a381a90474f4b7a7327819cd2675447

Observation 0755c884-03f5-45b1-9f02-8b01ec36648c · outbound

This paper cites Higgs audio v2: Text-audio foundation model.https://github.com/boson-ai/higgs-audio,.

Luna-TTS Family Technical Report Higgs audio v2: Text-audio foundation model.https://github.com/boson-ai/higgs-audio,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.550652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.204473Z digest=sha256:109268717daa8def8d699304355e755dd2853fff3b8f14a422fa3ce7394b1657

Observation d7044a85-bfc2-423a-bb0e-0912c153731b · outbound

This paper cites an unresolved cited work.

Luna-TTS Family Technical Report Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.215711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.215711Z digest=sha256:73abf0514cd006317b667c3a90d79ca17ad9152b0044db028f57d7ed7562a25f

Observation a9d5a483-a2bb-4c5e-ac72-000429ec4f77 · outbound

This paper cites d1: Scaling reasoning in diffusion large language models via reinforcement learning.

Luna-TTS Family Technical Report d1: Scaling reasoning in diffusion large language models via reinforcement learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.491484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.220764Z digest=sha256:de56b41f46bf4871fe9c5881bc8191fb94cb669c74dd869c8665235bb98cc593

Observation 7c172714-8c34-47fd-92c7-c7b975bc6b45 · outbound

This paper cites Training dif- fusion models with reinforcement learning.

Luna-TTS Family Technical Report Training dif- fusion models with reinforcement learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.471220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.226146Z digest=sha256:0dd5ab2b417d5878597d7d38f23841682a2b0c36c0a5818c7a10400fe98f9e03

Observation 16a08523-06b5-4c92-9199-b4ca8ba1e95a · outbound

This paper cites Mask-Aware Policy Gradients for Diffusion Language Models.

Luna-TTS Family Technical Report Mask-Aware Policy Gradients for Diffusion Language Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:40:25.729606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.231908Z digest=sha256:a9cc4e502be2c09f7a91fc4fddb8455bb1c03fd3a97414d9f92939ea4141796c

Observation a7022029-a0c2-4541-bc6a-ee6b96531f1c · outbound

This paper cites Classifier-Free Diffusion Guidance.

Luna-TTS Family Technical Report Classifier-Free Diffusion Guidance

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.237713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.237713Z digest=sha256:ef1d27c6e060371550bb3ac277cea61f165979858d2acb1c735e68bc32a95f21

Observation 0999b830-8486-4a30-aef6-9956eb221776 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Luna-TTS Family Technical Report Proximal Policy Optimization Algorithms

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.245620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.245620Z digest=sha256:f68e61ff69500a8a9c4d1f6d0a7534e1cb42242d2d829538b0d742f076a977d1

Observation 72562f2b-97a4-47bd-9f3e-7245ec759b31 · outbound

This paper cites vLLM-Omni: Fully disaggregated serving for any-to-any multimodal models.arXiv preprint arXiv:2602.02204, 2026.

Luna-TTS Family Technical Report vLLM-Omni: Fully disaggregated serving for any-to-any multimodal models.arXiv preprint arXiv:2602.02204, 2026

Reference 64

Resolution
verified exact
doi, observed 2026-08-16T00:40:25.663788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.251439Z digest=sha256:04e37ff71281f810db0958bad9fd02d4f4750e3d4c832006a786664ed28d3a02

Observation b3cb0fd2-0409-4b35-a793-3ff3247fa28c · outbound

This paper cites Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling.

Luna-TTS Family Technical Report Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.256850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.256850Z digest=sha256:e0edae50b795978cc2a3df453b5f8b447de38d56aaeab8006ed6daec68a6ccce

Observation e9b42b5c-8d5f-42b7-b6cc-28bd80a5fdf0 · outbound

This paper cites ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching.

Luna-TTS Family Technical Report ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.263064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.263064Z digest=sha256:e68abfdadc0998d99f5b866fd8ef3d158393cc5d5d68bb6c1f5e7b5153561d12

Observation 3383c818-eeea-462f-92cd-7b1c7f92a210 · outbound

This paper cites V oxCPM2: Production deployment and inference performance.https://github.com/OpenBMB/ VoxCPM, 2026.

Luna-TTS Family Technical Report V oxCPM2: Production deployment and inference performance.https://github.com/OpenBMB/ VoxCPM, 2026

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.451743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.268599Z digest=sha256:07d30414c4742592d5232ba073532488ae10c70208a44c6ff8708211d32a8c11

Observation d7d53a26-e462-4c9b-8a31-43f6f3272ed7 · outbound

This paper cites Spark-TTS: Nvidia triton inference serving.https://github.com/SparkAudio/Spark-TTS,.

Luna-TTS Family Technical Report Spark-TTS: Nvidia triton inference serving.https://github.com/SparkAudio/Spark-TTS,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.433402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.274343Z digest=sha256:fd5699e9d2f11d5ad87c2d531414e489589844f53c102719d70e21f6b498d2e0

Observation 14636bfa-2c68-48a3-859d-dde1e4b205a2 · outbound

This paper cites SGLang-Omni: High-performance multi-stage pipeline framework for omni models.https: //github.com/sgl-project/sglang-omni, 2026.

Luna-TTS Family Technical Report SGLang-Omni: High-performance multi-stage pipeline framework for omni models.https: //github.com/sgl-project/sglang-omni, 2026

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.393547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.285545Z digest=sha256:038615fe7021ab744e602fb88897933c6ee9c03056c0e3b95ca44b0cd19f438e

Observation 6b387dc5-d8fc-4366-a29e-eb98ee7771d1 · outbound

This paper cites an unresolved cited work.

Luna-TTS Family Technical Report Unresolved cited work

Reference 70

Resolution
parse uncertain
raw_fallback, observed 2026-08-16T00:40:27.413198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.279969Z digest=sha256:10406f8a2361ea0af6f08704c1f6a0d8d620dfe45d032d79514326aee717cf4f

Observation ef6be5de-06d9-48ce-a62c-9e32788f4257 · outbound

This paper cites Models: Eleven flash v2.5.https://elevenlabs.io/docs/overview/models, 2026.

Luna-TTS Family Technical Report Models: Eleven flash v2.5.https://elevenlabs.io/docs/overview/models, 2026

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.353717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.298092Z digest=sha256:6ba5a5c688ce5f6ce0e33b74710a820a952469dc93f9ead4d5526e99fe24a83b

Observation ef907eb4-eeed-4b26-914e-d23be500c154 · outbound

This paper cites Sonic 3.5 self-hosted hardware selection and latency.https://docs.cartesia.ai/self-hosted/ hardware-selection, 2026.

Luna-TTS Family Technical Report Sonic 3.5 self-hosted hardware selection and latency.https://docs.cartesia.ai/self-hosted/ hardware-selection, 2026

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.374753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.292032Z digest=sha256:b2227ea3569000dd8943fd734a5429f79847b42bf40faacc2a37db433940641e

Observation 6a673219-f2fd-4fbb-8ab4-cfdbd7c31cff · outbound

This paper cites Text-to-speech models: Play 3.0 mini.https://docs.play.ht/reference/models, 2026.

Luna-TTS Family Technical Report Text-to-speech models: Play 3.0 mini.https://docs.play.ht/reference/models, 2026

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.316937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.311830Z digest=sha256:3cac6594b1e82af80127893e530e225eefeb99485ee7c358001284b0d6456912

Observation 7acea1b9-de6a-498d-98b7-e44b577c018d · outbound

This paper cites Text-to-speech: Octave 2.https://dev.hume.ai/docs/text-to-speech-tts/overview, 2026.

Luna-TTS Family Technical Report Text-to-speech: Octave 2.https://dev.hume.ai/docs/text-to-speech-tts/overview, 2026

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.335785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.304257Z digest=sha256:a37f63bb4af7e0e43afea9c7a551bae5f6de8568ebc1125ff1ef7136b7dd533b

Observation 68e7dfd8-0b84-44cb-8346-d900beb3e6e8 · outbound

This paper cites Vibevoice-realtime: Real-time streaming text-to-speech.https://github.com/microsoft/ VibeVoice, 2025.

Luna-TTS Family Technical Report Vibevoice-realtime: Real-time streaming text-to-speech.https://github.com/microsoft/ VibeVoice, 2025

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.275358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.323406Z digest=sha256:82ccb2237316885902908488320f882644cf6215b7b85a615e0d31220e2d2dc0

Observation e9c7d253-39f1-4494-855e-a0960c789a68 · outbound

This paper cites Aura-2 text-to-speech performance.https://developers.deepgram.com/changelog/2025/5/ 14, 2025.

Luna-TTS Family Technical Report Aura-2 text-to-speech performance.https://developers.deepgram.com/changelog/2025/5/ 14, 2025

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.296430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.317697Z digest=sha256:630b453e95d1c9eefd887712c1f4b1d5924f7517662583327c32ab01ddebb1b1

Observation 00097372-654a-40da-b6a6-9e9f4c095d87 · outbound

This paper cites MiniMax Speech-2.8.https://www.minimax-speech.com/, 2026.

Luna-TTS Family Technical Report MiniMax Speech-2.8.https://www.minimax-speech.com/, 2026

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.254329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.333666Z digest=sha256:488e3facfe89b2f415eaae3bc599dcd6a824aa9bda6ff5dc556b61eef22b7c48

Observation 1ae9d38f-119e-42f2-afe5-0c9ec74fcc97 · outbound

This paper cites VoxCPM2 Technical Report.

Luna-TTS Family Technical Report VoxCPM2 Technical Report

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.328324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.328324Z digest=sha256:b3a155dbbbdf868baefade95bb0b3db937d13eaf31f5cf6a4e449bb6b13de762

Observation 60999807-8c61-4f19-8227-1ddfcbf47fa4 · outbound

This paper cites Introducing S2.1 Pro: Our most expressive TTS model yet.https://fish.audio/blog/ s2-1-pro-free-api/, 2026.

Luna-TTS Family Technical Report Introducing S2.1 Pro: Our most expressive TTS model yet.https://fish.audio/blog/ s2-1-pro-free-api/, 2026

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.207515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.343354Z digest=sha256:3c6a84b4236dae935a1fb45c0915190ce7395b4430b6c096dd20dee7ca71d5c2

Observation f7499e1e-025f-4953-b3d0-1b59187a9342 · outbound

This paper cites Eleven v3.https://elevenlabs.io/v3, 2026.

Luna-TTS Family Technical Report Eleven v3.https://elevenlabs.io/v3, 2026

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.232812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.338581Z digest=sha256:4d0fa1e6194041f1da02f2d0c7abfb57f2a8e096d977ee4e208110a01d0bdb01

Observation 8ae3aca2-7477-40ff-bb5e-fe526ba8a6d2 · outbound

This paper cites NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations.

Luna-TTS Family Technical Report NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.353780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.353780Z digest=sha256:5119bf7a7a795d6329fba43b1be5c067549128e9cd377aca24f21fa66fef20cc

Observation 222b54e8-2444-4ee5-8629-f2ba24accbf7 · outbound

This paper cites NV-Bench: Benchmark of nonverbal vocalization synthesis for expressive text-to-speech generation.arXiv preprint arXiv:2603.15352, 2026.

Luna-TTS Family Technical Report NV-Bench: Benchmark of nonverbal vocalization synthesis for expressive text-to-speech generation.arXiv preprint arXiv:2603.15352, 2026

Reference 82

Resolution
verified exact
doi, observed 2026-08-16T00:40:25.552584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.348470Z digest=sha256:e0bbba833ba7df3433885a39c711c2507ab237a97ab770e688339f10c7d56b9b

Observation 595aedc9-407e-4208-9e46-087b510c01d4 · outbound

This paper cites Emotional voice conversion: Theory, databases and ESD.

Luna-TTS Family Technical Report Emotional voice conversion: Theory, databases and ESD

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.364583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.364583Z digest=sha256:dfa74ce62ec6c5cc144e861bb640a5965ac49f0d3b56317bbfb451b15f069e89

Observation 79e1c23a-6dec-4b28-a9fc-31e0ec34b4d5 · outbound

This paper cites Gemini 3.1 Pro model card.https://deepmind.google/models/model-cards/ gemini-3-1-pro/, 2026.

Luna-TTS Family Technical Report Gemini 3.1 Pro model card.https://deepmind.google/models/model-cards/ gemini-3-1-pro/, 2026

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:40:27.179821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.359507Z digest=sha256:999c500ed8886b29100078049c9d516c628ac481a109c18e887c43fcbc0bea2b

Observation 1cf04fc3-8cda-4216-8d6e-27e1447869d4 · outbound

This paper cites an unresolved cited work.

Luna-TTS Family Technical Report Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:40:27.157862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.375381Z digest=sha256:2c5a58deb30c4648e7222d54bab895203cddd6b83cd2ccea44d1e44a268a0764

Observation 5d6aec85-664d-4bec-924d-c85b29137d5b · outbound

This paper cites emotion2vec: Self-supervised pre-training for speech emotion representation.

Luna-TTS Family Technical Report emotion2vec: Self-supervised pre-training for speech emotion representation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-16T00:40:25.369872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:40:25.369872Z digest=sha256:02b5266df6834ce4824bcc7d6cb7fdd6e521e36b46ae345c4fabc1a75b2567af

Observation 96d3bf61-c75d-4e0c-8bc0-956f843c7efb · outbound

This paper cites MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech.

Luna-TTS Family Technical Report MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:40:25.756497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.380445Z digest=sha256:bb4442287babc1daf0c1c69f546a5bfc8c9129fabad786ec3e89f3854d2af1d4

Observation 31f381b5-43a6-427e-b071-63a78a11a2c6 · outbound

This paper cites an unresolved cited work.

Luna-TTS Family Technical Report Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:40:27.529429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:40:25.210326Z digest=sha256:89b86428c7fb593b419285bcf16bd4488a6cc0552a3b77e7889515822039ceb4

Pith citing papers

No inbound Pith citation observations are available.