Pith. sign in

Paper Citation Record · LEDGER

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

As of 7 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 100 inbound Pith citation observations for arXiv:2406.02430.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.02430 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T12:26:37.300599Z

measured 145 of 145 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 100 of 131 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:23:15.641196Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact35
  • verified fuzzy7
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

5
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 4756c9b5-0052-4463-b6c7-7f5322bf9b4c · outbound

This paper cites Streaming voice conversion via intermediate bottleneck features and non-streaming teacher guidance.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Streaming voice conversion via intermediate bottleneck features and non-streaming teacher guidance

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:26:37.509092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:7bcccaeda9fd4e0d7d689f20ebd4a216c3bb1c7a9feffbee68b0075dcfc72c61

Observation 7a3c3035-377d-4058-b211-d959de06cbb1 · outbound

This paper cites StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.331901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:2087a47f4d49e052a3f5bf3c3930ec85a833b6448a1427ecaa5e30fa894737ec

Observation e2f60dc4-b155-4d77-ad62-d3f221336643 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.337799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:e8b965a333db46655c8a11f68d4032a2578d3220d743296884f8c919e52159cc

Observation 53c5908d-2c1c-4b39-8fe2-588f9940ae35 · outbound

This paper cites Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.342857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:9f9a8e535fb67032f9e2020f5c2daecd2ed4886e0bbf10677df865ec568aabbe

Observation 006c7288-10a0-4add-ba1c-330ff6c2651a · outbound

This paper cites Deep Reinforcement Learning: An Overview.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Deep Reinforcement Learning: An Overview

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:26:37.347675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:382ab5903ac49c9e36d5217859a8ef1143264c0acb9de604c396cdb8f136e504

Observation 979bebc8-0d3e-45fd-8ca7-22bb17eac15d · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:26:37.352136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:5981ee1c06a07425df44c3ca621679b46bd33952162cb93b6f0ea408c1fb033b

Observation 644271db-4c5d-4826-9649-6a0f66576348 · outbound

This paper cites ResGrad: Residual Denoising Diffusion Probabilistic Models for Text to Speech.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models ResGrad: Residual Denoising Diffusion Probabilistic Models for Text to Speech

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.358488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:0ee8a64cbb4850eca3624366720bb129e756fe51a54e14df326e5aa6695cb4d2

Observation 1d4b87f3-9a05-4801-bdb4-14a32e18b6dc · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models LLaMA: Open and Efficient Foundation Language Models

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T12:26:37.364028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:01b4ab37fd9d50337e2bd646170bca42f40a2f6d8ee6589e07557fcdbeb6cd91

Observation 36c79e26-f01d-4b1e-954a-a6c8070fc732 · outbound

This paper cites Better speech synthesis through scaling.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Better speech synthesis through scaling

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.369560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:9605f42bd905d9f2371189b643e52a4354e91e9e194eb88fe65f613051ef019b

Observation d23726c2-113b-44f4-aa51-28e4166d5cf8 · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.374042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:3763796ba4e5a5b8ab349eb03042be63a0fb4df2fb026e2f009ad4a086825df2

Observation 0441f097-3029-4902-b262-3fe63dad3c05 · outbound

This paper cites Glow-WaveGAN: Learning Speech Representations from GAN-based Variational Auto-Encoder For High Fidelity Flow-based Speech Synthesis.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Glow-WaveGAN: Learning Speech Representations from GAN-based Variational Auto-Encoder For High Fidelity Flow-based Speech Synthesis

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.325985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:866ffb8b7d37aaba258591a3b811b910ade2ae855f7f64cf093d57f4320228e1

Observation dd3e66ab-0ea7-41fc-af35-223fa74ac991 · outbound

This paper cites Basis-MelGAN: Efficient Neural Vocoder Based on Audio Decomposition.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Basis-MelGAN: Efficient Neural Vocoder Based on Audio Decomposition

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.417345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:1d6a87aa6e4182d2e4fceb3d0a22bd83fee07bd967d0ef3afb790935291c2f49

Observation d7ebdee7-0e41-4c7c-89d6-d2cb89f158a1 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Common Voice: A Massively-Multilingual Speech Corpus

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:26:37.422301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:3e81daf1ce5401a0b15df4813e28703972b4cd3f7eace1134f8cc13ab9266f69

Observation 8e025bd0-40a1-40b1-89aa-44c5492b1da1 · outbound

This paper cites DiDiSpeech: A large scale mandarin speech corpus.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models DiDiSpeech: A large scale mandarin speech corpus

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:26:37.518068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:2af0a0e80c19eb32516399b4633c68130e6933e456ca58b483b9a50af4f3a9c4

Observation 029a82d4-079b-478f-abb3-40d13ae9a8f8 · outbound

This paper cites FunASR: A Fundamental End-to-End Speech Recognition Toolkit.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.427555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:b2fb8f6800ce973f5128ace1459c14fafa89c89250f7e313657b2677b7cc009a

Observation e3ced301-f751-4fa9-a4b5-f08d66b409ff · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.432620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:27f414abea1d826e14ed1f18147957550785b847c994d1bb6904505b6be62dc3

Observation f7720dba-20cb-429d-9c36-7acfce416ee8 · outbound

This paper cites Controllable and Lossless Non-Autoregressive End-to-End Text-to-Speech.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Controllable and Lossless Non-Autoregressive End-to-End Text-to-Speech

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.437031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:3e17f4e952204f9f9149127b00cdf2ab03a0ddd1e5e1be0e264eeafde43af093

Observation 4888ea06-1982-4030-9480-89c7d4842d21 · outbound

This paper cites LibriSpeech: an ASR corpus based on public domain audio books.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models LibriSpeech: an ASR corpus based on public domain audio books

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:26:37.512281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:6177ddc30d18ac8516a310b69f129da16e079f7b40418fb6c8a090c1f10b7388

Observation 3bf8aecc-c6a2-43e1-bde2-2c0212de1e5e · outbound

This paper cites WeNet 2.0: More Productive End-to-End Speech Recognition Toolkit.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models WeNet 2.0: More Productive End-to-End Speech Recognition Toolkit

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.442050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:3489e402fc59b4a15fdcbd8ab9f527e92c12f38b0ec6cb736285822594e959a1

Observation 244c1ec1-8efe-40cb-824b-9ec48aa3bb21 · outbound

This paper cites Developing far-field speaker system via teacher-student learning.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Developing far-field speaker system via teacher-student learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:26:37.520505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:d834f865c48b8d421df268d0a84a7dde6e5a3c2a2cb6ba1a2c3acba161b8457b

Observation be268926-e6b8-4033-9530-b496693524b1 · outbound

This paper cites VoxCeleb: a large-scale speaker identification dataset.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models VoxCeleb: a large-scale speaker identification dataset

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.447207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:97e244a6c59929d1df2fe11cf9b01f3e8c44eea56387e651826f836ee080ce8e

Observation feb995b5-9758-4d62-b15d-8bf198e9a477 · outbound

This paper cites Singing Voice Synthesis Using Deep Autoregressive Neural Networks for Acoustic Modeling.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Singing Voice Synthesis Using Deep Autoregressive Neural Networks for Acoustic Modeling

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:26:37.451102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:ad70ac4d8e34987b9728940e01c81b9cea97ea83fdf99f1fc262a65f3242ef56

Observation 1a553839-b497-4fc3-8507-673b51d844c1 · outbound

This paper cites LiteSing: Towards fast, lightweight and expressive singing voice synthesis.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models LiteSing: Towards fast, lightweight and expressive singing voice synthesis

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:26:37.515667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:477e727541957f90a04776cfa47cd4544585d3fca353a8fa61aaef284d621188

Observation eaab47d9-a256-4d2c-bb7c-82f7902d3886 · outbound

This paper cites Prosody-aware SpeechT5 for expressive neural TTS.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Prosody-aware SpeechT5 for expressive neural TTS

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:26:37.522911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:a6a33bca6a2fefb07b4a71fb97c5aa1ed888bf4bdd5a440c0e6c4c9ec5205c0e

Observation 8d25e85c-fb2c-4627-a062-0503dff6d454 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:26:37.455658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:b410b2dee0393e128432ebbf960bc11b651abf71981420d1a11578cc263fdcca

Observation 75228981-875d-439e-84e7-34cba46e75a8 · outbound

This paper cites Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:26:37.460225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:9bb5ed3f391c2168682ebd78ce009170c8b9dc3e74ccf8d31eb861cdfe841c78

Observation 5126f41a-d88e-41e7-b8cc-29a36906a6a9 · outbound

This paper cites AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:26:37.464522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:b8ec5ec41453fb00ba2678ea2784a3beee198e95d455748e9766ac7b0b212341

Observation 2416902b-a921-482c-bde9-5c4c5826eb96 · outbound

This paper cites Consistency Models.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Consistency Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:26:37.468540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:be41d115cc86cb27007885ec27501c7010adec6903b7129c4895edab37d8d781

Observation f8a34ae0-e1a1-436d-acc6-d9a3c263153c · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:26:37.472833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:7225fdde47dde16fd7e0f909a0b010a63f502fa20e29c4ab660edd45f0ad1806

Observation 8fa63e6b-41c5-4885-a1e7-7d0509da4b5a · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:26:37.476810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:b1d5d7fc5597a5cd1f97e25c30268f18025697bb92c08f4f5b2eb6ffc58c9776

Observation 8e4d6af0-fd01-47db-b303-3273d6c9719c · outbound

This paper cites A White Paper on Neural Network Quantization.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models A White Paper on Neural Network Quantization

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:10:55.058388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:f75844281fea22417c8e2ba667ed326d406029946a761ca68a44ac4187f5e3f9

Observation 2435136e-38f8-4e99-84cf-e3400f2314c3 · outbound

This paper cites decoupleQ: Towards 2-bit Post-Training Uniform Quantization via decoupling Parameters into Integer and Floating Points.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models decoupleQ: Towards 2-bit Post-Training Uniform Quantization via decoupling Parameters into Integer and Floating Points

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.485178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:f92cf5dbe346a5c10f04f7ec817193e840aa1f093846b76e76b178a7efb3685a

Observation 6092f1ee-f92a-4a94-96ec-7c9a045bf1c2 · outbound

This paper cites VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.489334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:63f25e18b2a15df6220ea43d6ca4300b219bacf80ad8e71840b672ebbe7bd851

Observation 1b46cae5-38b9-4e85-b7a8-66bdc50c31c2 · outbound

This paper cites HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models HierSpeech++: Bridging the Gap between Semantic and Acoustic Representation of Speech by Hierarchical Variational Inference for Zero-shot Speech Synthesis

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.493296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:303752f7b969289eaa3035817d1bb20f40314be05e8c108b3596227d53d4dce3

Observation 09a8622b-3123-44d5-bc89-b24fd7ab793b · outbound

This paper cites Zero-Shot Accent Conversion using Pseudo Siamese Disentanglement Network.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Zero-Shot Accent Conversion using Pseudo Siamese Disentanglement Network

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.497056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:9dce6da6ba4e40dc42140c17525e39afc685e1f59f11ea4031e90594c471e669

Observation c7ee5465-f4e3-43ff-a7f0-10f80a53ac2a · outbound

This paper cites Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Diffusion-Based Voice Conversion with Fast Maximum Likelihood Sampling Scheme

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.501347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:fc6f05594e7c4d36f0a99623d3d0b05badc71ea1d842531417983674796d0243

Observation 5e45a6f5-b77a-4f83-867a-09990480a4b3 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Proximal Policy Optimization Algorithms

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:26:37.505827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:fd70d482aee268a4286e4d75be4c416409fbfb655ec81d2fe7cdf86326a0b76f

Observation 985d286e-bda1-4c57-81d3-b0dcdb611313 · outbound

This paper cites Diffusion Model Alignment Using Direct Preference Optimization.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Diffusion Model Alignment Using Direct Preference Optimization

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.380158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:09a1e826609a8854c8fe63da985d4ef6e0e89c4ea3136c9922fa26746ba2fb37

Observation 2cae3087-dbe4-41ee-a48a-66159963c655 · outbound

This paper cites MusicRL: Aligning Music Generation to Human Preferences.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models MusicRL: Aligning Music Generation to Human Preferences

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.384884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:dead189f59eb40cb889ac7c9cfd74394499ebdc8ca473b42e2f5cdf46539e175

Observation 2b1927a3-db00-49e5-9f38-c32ea6a04dea · outbound

This paper cites SpeechAlign: Aligning Speech Generation to Human Preferences.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models SpeechAlign: Aligning Speech Generation to Human Preferences

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.389182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:b5e1577d75de5c8451faca94fbe552abde64e0d270e19c5fdd0848b06c79bb83

Observation aa06ca70-0c3d-4002-80f4-1a3e4a9dc51a · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:26:37.394037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:a98dc7899ed77c7755cdb5a8fbd27a38c813684500a9500b90eefc0fbbdfda7a

Observation b1a582d4-f2e3-43b8-a5b4-83a6915af56d · outbound

This paper cites Minimum word error rate training for attention-based sequence-to-sequence models.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Minimum word error rate training for attention-based sequence-to-sequence models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T12:26:37.525092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:6257798a0275295b7ccaf813a694c86bbd24ce5db3946ccd7bcacaa262950a8f

Observation 3fffca85-55af-4cda-a00a-bed68e84db18 · outbound

This paper cites Transforming and Combining Rewards for Aligning Large Language Models.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Transforming and Combining Rewards for Aligning Large Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.399932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:e039d38fcfdb3496d61b9f1d6f99c2681c50f11221d957bc00e37171b61d2f1f

Observation 8269caef-9368-47f3-809b-a6a66bde15fd · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-15T12:26:37.406253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:50038c3a04de9b7ea1ce77a4671b438a209ce7dc147ebe6b3a600443d19aa6d4

Observation cb7da3c8-9b5c-419a-aa7b-70d4dd5b084d · outbound

This paper cites SpeechX: Neural Codec Language Model as a Versatile Speech Transformer.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models SpeechX: Neural Codec Language Model as a Versatile Speech Transformer

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.412512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:25f1e6a164724111e05942f423816e11c4d39c497167b8b6ed7d00d788523efb

Pith citing papers

Observation 041024c9-5f47-488e-9eb4-24126add3958 · inbound

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS cites this paper.

Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-23T20:33:25.586949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T20:31:24.494060Z digest=sha256:2674d54274d2fc047a4576a2ad7cb48ff7c8c6d5f3d15156cb7272aeccd51315

Observation 6000b29f-4584-435c-9497-184926710ea2 · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:06:41.513885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:46d8f2926295bc8559a93d36619b4a52392a27a52d62748a50f06f7ad760b654

Observation 2287c955-3603-42e4-99a2-762e4db69d33 · inbound

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models cites this paper.

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T06:19:09.507440Z digest=sha256:e8a3e614561a0e52f135d6373852fe08c3d11f0b63e8b067bbdf6142ed54742d

Observation 0024d1c8-e6d0-4bcc-9ca5-e0bd5e3fc405 · inbound

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction cites this paper.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:39:48.408958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:e50b2ce4ef38d619f7cfd5ca12853cad746768765c83e25aa96cbccc88ed9acf

Observation e91e50a7-28a0-41b1-a4ec-af7cc51c30ab · inbound

Qwen2.5-Omni Technical Report cites this paper.

Qwen2.5-Omni Technical Report Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:54:03.225439Z digest=sha256:55c68fc1eb1ceb841e4b5101e37f7aa4e868b1ccf40104d6fba667e184dd9b7d

Observation 66210f73-2ac1-4062-9482-daa2511608c5 · inbound

SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement cites this paper.

SeamlessEdit: Background Noise Aware Zero-Shot Speech Editing with in-Context Enhancement Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-22T14:54:54.683917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T14:53:09.885971Z digest=sha256:d46494dbf7812a04f15354485db678ef7260ed3e31ce7e6bd540b7474f2fcee9

Observation ca8bfa3e-ec9a-44ad-b667-1584f0572553 · inbound

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding cites this paper.

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:15.641196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:15.641196Z digest=sha256:9884a58b4b189b21a413d98651224cf71e86cf04d5f8e7135e7daae48627f999

Observation 9d33002b-6b3c-47b2-89c1-3fcdaa12fcef · inbound

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information cites this paper.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:14.619002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:14.619002Z digest=sha256:b66fe4f39345eadf35e5e4bcb3246cc54299ae850d9babbdd807d90460752d78

Observation c0938d39-b2e4-4570-afee-d5c45c76504f · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:27:25.538207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:6cd1f09af7415d2c92e5a19a24d9ff98077cfad8c8080b481705911894599511

Observation 8d541a03-a710-477c-a4db-f2ce42905748 · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:57.780080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:57.780080Z digest=sha256:b979749bae24b029ac473625a211462cb873f72a90aa9a522445e64d1916d0bc

Observation 1d4d6018-6c07-457a-8bcd-920d8e4b9f3c · inbound

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment cites this paper.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:53.546922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:53.546922Z digest=sha256:24270923d173580aa8325afe4731f974d12c8adf77c2b96053e863147cd15dc5

Observation 8d788c02-3a57-4d3a-9989-b5a4490215d9 · inbound

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling cites this paper.

Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:09:18.881600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:09:18.881600Z digest=sha256:07eaa48ca05f7a89e05514e9f45302f7bf3fe619c12dff846775072b15179802

Observation 061b2b2b-82f4-4c5a-84d2-8a4e4e9d4b26 · inbound

Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages cites this paper.

Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:54:23.772789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:54:23.772789Z digest=sha256:0b4b2d26a34649b4a5f39954b4447b7db156949f702636497b089decfc45800b

Observation 914a2250-af2d-4a02-841e-30a570b8fdb0 · inbound

MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation cites this paper.

MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:27.861952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:27.861952Z digest=sha256:918ed4b9e75e4f50b42372a1e2cf6bce1210f9b2550f69ba48abc2323d04dbc3

Observation 6c068f11-9a6e-4a59-b3d1-ebd2798c923b · inbound

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction cites this paper.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:32.751404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:32.751404Z digest=sha256:6e4a7d35ea2636214e5516e2136bff51f4b037bdd1a5d22ec5f8967e7b9679a3

Observation bd8141c1-7133-4193-a7d0-aac7c6faba2a · inbound

RoboEgo System Card: An Omnimodal Model with Native Full Duplexity cites this paper.

RoboEgo System Card: An Omnimodal Model with Native Full Duplexity Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:37:29.364137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:37:29.364137Z digest=sha256:b300c509476b920e16b6da6772295c1e247821f8a8c995081ed1dfa27e007be8

Observation 97c7ba52-f118-4652-bb40-c2f3d27a5a33 · inbound

Ming-Omni: A Unified Multimodal Model for Perception and Generation cites this paper.

Ming-Omni: A Unified Multimodal Model for Perception and Generation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:58:06.798089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:58:06.798089Z digest=sha256:abefabfab6092af9e2b11083fb5f0e9ee9a206815387c26fb87951674936f785

Observation 42ac2bef-7eda-46a8-93e0-beeed2e49c55 · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:43.623663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:43.623663Z digest=sha256:64b8bbb2c1475ff571856d1d8aa251a5bb0400d50ef9efcbd73860e2d8d86f35

Observation 811b01f7-0773-4610-a07c-7bf31ea74375 · inbound

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching cites this paper.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.675822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.675822Z digest=sha256:8b938ccd30c4510e2c498ba1113143d2fbea9fa3124031e237e6efcfb3a2fcda

Observation d0d2d73e-b959-474c-abb0-a24333169bfc · inbound

SimuPanel: A Novel Immersive Multi-Agent System to Simulate Interactive Expert Panel Discussion cites this paper.

SimuPanel: A Novel Immersive Multi-Agent System to Simulate Interactive Expert Panel Discussion Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:43.838451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:48:43.838451Z digest=sha256:2248faabd9b2853f7eb15ae18e172b40d1d4e170ee888c38ab76441dc4eb9053

Observation ba6482a5-3d62-479f-b595-93150211c630 · inbound

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech cites this paper.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:51.474748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:51.474748Z digest=sha256:f0e0baf1064d18ebdd8b189087cf89c7d1e44f77b326b68e14637174c2b185c2

Observation 3fff0af9-2b2b-4938-aa93-41cff71acb69 · inbound

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy cites this paper.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:00.703156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:00.703156Z digest=sha256:b090d2a988b8cca15817f3cc0460b1a404a27f680607a61965ae8cff5925bb94

Observation 0b63ce8a-c9bf-4023-b1e3-db6397667678 · inbound

Differentiable Reward Optimization for LLM based TTS system cites this paper.

Differentiable Reward Optimization for LLM based TTS system Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:03.946996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:03.946996Z digest=sha256:870cf86bab53ee5827d23d35950499eb765434b264ff276cdc1e51fbdcbc0fcb

Observation ae0691d8-4a74-473f-aada-22d1c6418655 · inbound

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models cites this paper.

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T03:42:44.523919Z digest=sha256:be0c3254a9779f5f90630de0a70aa4504d463d9914fe1c143c48de2c6170fc00

Observation c215fa50-64f0-4d60-b724-468a474cd059 · inbound

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching cites this paper.

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-19T04:32:03.700577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T04:29:41.285194Z digest=sha256:c47e58b2325ecf52c87ed94580f08ef9dd80a1426b085cef982d3ad04e8786c2

Observation 971e9b2c-bccb-40d3-ac1b-bd6a2a5cb146 · inbound

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations cites this paper.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.452658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.452658Z digest=sha256:f29d30226e48e4ab74ba5a4e4d02f32d7fbd628218a7a1e38e2230970e6ee165

Observation 49409248-3195-4e3d-b17a-5e25c7bb2263 · inbound

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis cites this paper.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.884960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.884960Z digest=sha256:eac1d267963c864bd726843053021d3b9084941896a3c0994190a3340cb16854

Observation f313ab2b-dc9a-4109-a0a2-f5ea84ba8cf5 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:59:51.008000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:2c7a317a3f811f40e33447dbc19590317050e1e245a822ce1f0023098d8d0e80

Observation e0ae2f24-26f0-4756-b95f-8a1031fa1b01 · inbound

SplitMeanFlow: Interval Splitting Consistency in Few-Step Generative Modeling cites this paper.

SplitMeanFlow: Interval Splitting Consistency in Few-Step Generative Modeling Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:07:26.599209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:07:26.599209Z digest=sha256:7a1fb55d0a05aec82545b46b6efffaa4cba20924bf79a16bde79fcd5eb4aae07

Observation b4245843-32f7-45af-b82e-a5839e43af1a · inbound

Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice cites this paper.

Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:52:42.942516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:52:42.942516Z digest=sha256:2e60b6c8415af3d5cfaf81bb7443160dbbc3a8e9b0402187c678da2531a279f4

Observation be7f1afa-59f8-436c-94e5-b0b99ba1b117 · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.208123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.208123Z digest=sha256:429d5ec95b603692df31ea431eed79343a162e3fe3fcbfb645df6d2540f3ce6c

Observation aed9d272-a437-42ff-bec8-c77046dd3603 · inbound

TTS-1 Technical Report cites this paper.

TTS-1 Technical Report Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.897698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.897698Z digest=sha256:46689f4aac1dd2fccd0ccfb489563dfd1849ba74feeeb881d706c7b8f048ed61

Observation 368df377-e60e-4603-b996-af6b64354af5 · inbound

Adaptive Duration Model for Text Speech Alignment cites this paper.

Adaptive Duration Model for Text Speech Alignment Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T11:34:08.681690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:34:08.681690Z digest=sha256:b8982d0ab953770006ed3fcb8f436fa35fca55a4ab703de7352a18ac7474c61a

Observation c770a0db-f760-470e-a2bc-9da33e913861 · inbound

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation cites this paper.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.065491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.065491Z digest=sha256:33a3a55bb9bb6fe1f4e96b0f400add8d12a8c4068c702a469f9078fc3cda3731

Observation 54bdc28c-fedc-4056-8ee1-8e20b8959378 · inbound

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation cites this paper.

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T23:43:03.170323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:43:03.170323Z digest=sha256:775ea92049c3b3eec8653af8e18a9c4dea4d202b230e0c2ac5f0ff7822614672

Observation 653e739d-550c-4d36-abff-8a836d6e1634 · inbound

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts cites this paper.

MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:04:50.143999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:04:50.143999Z digest=sha256:7e7d00e936c43654350062d69fb789ddcdec0782988614dbf41b3ce5847cfa35

Observation 65d4b235-9ce0-4f6d-bdf3-10d2af920dff · inbound

NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding cites this paper.

NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T18:39:27.824911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:39:27.824911Z digest=sha256:9143d35a1fc8f9ec85e3b8a2e66f9cc32d33ee0ed79ddc1e4131300badaf5917

Observation 958bcb74-b0d2-4b6d-b750-76a12a654377 · inbound

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis cites this paper.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.500071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.500071Z digest=sha256:a5f5256b5a21624b784c8f800aa872767787e04a4aa639a71ade7feea487ab8b

Observation f8d55cdd-b927-4d12-b893-ac9fee5a77e7 · inbound

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech cites this paper.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.204193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.204193Z digest=sha256:52e16aa106ade5b888f8945a86869c1e699c6b78617a80acb6d4f3ed5a5a3605

Observation 78740931-f52f-485a-baba-662d9142bc05 · inbound

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot cites this paper.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.422814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.422814Z digest=sha256:921b9e41bda288c91f00c64e5551f5c3cc90bbd6aa3b5f771ed468d57469bde4

Observation 413d3af7-a553-4508-a4ab-f335efe03ca0 · inbound

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration cites this paper.

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T19:11:59.476799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:11:59.476799Z digest=sha256:314fbf0a23c4c06ce2e37dbb788e59c5fa9529f8b47238f031bb026c14c45829

Observation 127fb2f9-b8ac-4151-a00b-fa8dde61b793 · inbound

Qwen3-Omni Technical Report cites this paper.

Qwen3-Omni Technical Report Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:20:37.406351Z digest=sha256:416bfe2c4b67b4127edb19bb3dffb51cdd9bf10aad15594b95631a0e26150813

Observation d7bd7831-89ca-4f66-824a-53913e906523 · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:01:24.425634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:23d67fc6cabfbe147b9f5a7f26fa3a576006fb1c3f8437844cfc2b97657fc8bc

Observation a9d33b68-1469-43b4-b45e-0df35f32b1cc · inbound

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation cites this paper.

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T01:03:15.285620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T01:03:15.183360Z digest=sha256:6f5dac89fb0a45d59cd1be5752e4efdd917e2eb9889191d49dedf539b5f06822

Observation ad9c0fc2-96f3-42d0-8681-5988cbee1899 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:29.658196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:29.658196Z digest=sha256:6561ad8cea5c1341d4582c83f88179a26ea65df286642c5e9bc49b372712072b

Observation 3ad9b056-fa54-43bd-9269-f8e379043065 · inbound

Two-Dimensional Quantization for Geometry-Aware Audio Coding cites this paper.

Two-Dimensional Quantization for Geometry-Aware Audio Coding Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-21T18:20:29.267476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T18:16:51.486807Z digest=sha256:f19ba4be8420b00d31d37a0edbee0bba52fed1c232d3b3cd7b1970023bb78f47

Observation 4e4ec9b8-24c5-477d-98a4-25da48619be5 · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:24:56.112886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:d5b86d334ba2531282ac79fdf63a7bc34a95f8f9b40b077e130bebe025cc3c46

Observation 762d14bd-1ad1-4767-9a1c-6ca2513c7d76 · inbound

Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards cites this paper.

Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T06:03:34.891329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:03:34.891329Z digest=sha256:601ec2927850d6bf858d8b27b06753ac1725e64b05e4c7c6caa404f5a92da215

Observation 2e099125-5d51-42d8-9179-f8b163ee6863 · inbound

Voxtral TTS cites this paper.

Voxtral TTS Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:515afb4f2f875a1e4d8a7604de57ed81fe102dbdfb8857d23fb1987623fea05e

Observation ddf0ac01-a325-4c9b-bd3d-9f2f54c30453 · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:308e3841359f4e67c97e6eecad70d5deaea38daf6968d8f04936ced3fa88af7a

Observation 344fe637-5357-4c17-9763-3bb213d23851 · inbound

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck cites this paper.

Controllable Singing Style Conversion with Boundary-Aware Information Bottleneck Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:51:38.030059Z digest=sha256:beb12684c8ed4f7ae1da6363e65c9e1d141f8b89724e1330ef5290a0bd66af2e

Observation 5923032c-c62d-4621-90d4-f0fac40f1c00 · inbound

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation cites this paper.

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T17:15:28.918204Z digest=sha256:181a495ac4e39218b972d5cf46192165391d4ef9d21bc0b153e5c28e468c0b35

Observation 93d3680e-622e-4165-aa15-bb465142f1f3 · inbound

WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models cites this paper.

WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T10:29:51.313835Z digest=sha256:44405a7db9af02e86d2b0d06ff385ceb7a3e8d34619716175013fa8d73dce923

Observation b45ebf96-f36d-4ff1-8d55-f2885248ae0b · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:04a6f692068c4c700c10905088ac1e599d02c61dbb71769221e1140003822f3e

Observation a6a1f961-7ef8-4cc2-ac74-94e3e8085089 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:ede6e156dd64197246e0409dc3111fda57eb1e0315badc69ba09a968708be804

Observation 368dbd12-4a7c-4adc-85a2-480ae8d16224 · inbound

Bridging What the Model Thinks and How It Speaks: Self-Aware Speech Language Models for Expressive Speech Generation cites this paper.

Bridging What the Model Thinks and How It Speaks: Self-Aware Speech Language Models for Expressive Speech Generation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:16:10.832592Z digest=sha256:3c9049e95c61e8df75cd16d79364be910c16f04875ca2c97a32fe5b915a25497

Observation 0babbef2-c98d-4d4c-b65d-27c612265560 · inbound

X-VC: Zero-shot Streaming Voice Conversion in Codec Space cites this paper.

X-VC: Zero-shot Streaming Voice Conversion in Codec Space Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T14:20:01.885544Z digest=sha256:478c109060db1206a2792f8404fbac2ccf782c6a0f37c6aa0deafc03e5ac4524

Observation 0f264bc0-690c-42dc-bfd4-a656e57854fb · inbound

From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench cites this paper.

From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:44:12.373082Z digest=sha256:e47a430c5205ea9476005ef95c6c71af266827fa4dfc793464f5224f3875f76b

Observation 5cd30102-480c-45e5-8003-39ec429f79ef · inbound

Qwen3.5-Omni Technical Report cites this paper.

Qwen3.5-Omni Technical Report Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T08:11:22.402552Z digest=sha256:530926c3ade78f871a87be8d281f5d151b65e7c8333bb2c290178ea1bcf44c92

Observation b3ef62b0-6daf-46df-a116-19ff86208f6b · inbound

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech cites this paper.

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T04:03:38.919545Z digest=sha256:aa46661085b6d9dd5277fa57a99fdcafec1bf878352fe46f1c598de51361a987

Observation bc83a086-21fb-4ef8-b518-a539e471edfb · inbound

Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation cites this paper.

Text-To-Speech with Chain-of-Details: modeling temporal dynamics in speech generation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T01:01:06.094276Z digest=sha256:ee31af7855c827c93d33bd13442f4fa7f58e0f0cccf601f5876c4f36fc5cc42d

Observation 1d6cf6c1-a7ea-4697-acb7-4c5191f789ed · inbound

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation cites this paper.

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T02:55:09.008954Z digest=sha256:403705f5907cea9511bd3f2f7c1b156bafb13bb8461ceb4148b19ac02dc40b6a

Observation 973c30f0-f12b-4bdd-b61f-d359795cec47 · inbound

JaiTTS: A Thai Voice Cloning Model cites this paper.

JaiTTS: A Thai Voice Cloning Model Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-07T06:34:45.357695Z digest=sha256:55572feefb7d75c3244e901e8f39f5731ff797f916285558575f8bb3cda05e7a

Observation f7a73220-847e-4795-90c8-943f60284c24 · inbound

JaiTTS: A Thai Voice Cloning Model cites this paper.

JaiTTS: A Thai Voice Cloning Model Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T03:09:24.541657Z digest=sha256:d62e7416b89b72c5d0905d0b2ece5ad2b2825df71d83151757a7413ccf723611

Observation 92f9b56f-4056-4c15-a7a8-528731586257 · inbound

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech cites this paper.

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:04:44.050889Z digest=sha256:f3e9113485a090aca91dad8033976ad272d22690e24f30ff03125a4565919c92

Observation beb68d03-d8a0-4a6b-83c4-21e8f1661d43 · inbound

Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic Cues cites this paper.

Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic Cues Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T18:37:42.783936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T18:36:34.351533Z digest=sha256:591b64764df45d12e3654fe5406d26902128bbc5bb8ce26c42bc26a0600bd37b

Observation 6c05834e-168f-4850-8991-eee0d38ab236 · inbound

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis cites this paper.

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-19T19:02:43.586091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T18:58:15.288299Z digest=sha256:49c08ed38a14483382dc2bc12a8d64868dddee42edfe2a3cebfe3ada87827d52

Observation 59e97aa2-3ac4-4d1f-b1a7-bde25cad6181 · inbound

Taming Audio VAEs via Target-KL Regularization cites this paper.

Taming Audio VAEs via Target-KL Regularization Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:53:23.190532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:b75bd4731c5739bcd95ce69538064c9b2bcbdb24902148f6cb6277693f080f03

Observation 52ccb11b-2904-498b-b600-24adc1f572da · inbound

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech cites this paper.

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-21T02:33:55.429131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T02:32:26.122526Z digest=sha256:17ee29357a09d8bcdf51cae5a43f3660e56df35b354e2db126076842596e43cb

Observation 283516ce-deb9-4eb0-92d9-253a24097c53 · inbound

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cites this paper.

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-22T03:00:58.966721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T02:56:06.910758Z digest=sha256:32502ad0d045a893767c7f6d003ed0a9248f980d1b1ba79f91846bdd7a6db9fb

Observation 205d5100-80f6-4707-acfb-43c77163227e · inbound

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching cites this paper.

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T13:29:53.155695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:29:53.155695Z digest=sha256:4f14343950272f4b1514535ba061cdc763d00d2308a1b5aa0a869ba9f24ade22

Observation 3db25c6d-0be8-49f4-a71a-530a3de1d0b9 · inbound

Raon-Speech Technical Report cites this paper.

Raon-Speech Technical Report Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T08:22:44.347941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T08:22:44.347941Z digest=sha256:d35d4a3671be1c61522cab1d1ba869373c2b1bcb421bc8884c16fbffb7de2f05

Observation a795d270-d0da-4fde-a878-3e70b2b63207 · inbound

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis cites this paper.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-29T16:23:39.877202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:85ea6ee89656e928e96da3d33cbefdd55e234649204f871efe7f19fb8489a9ac

Observation e85a2a67-8566-44eb-ba9a-0b7ffd97609e · inbound

Native Audio-Visual Alignment for Generation cites this paper.

Native Audio-Visual Alignment for Generation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:03:14.765078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T07:53:59.795431Z digest=sha256:6f1cccee708aab23766787d895acb1d1cad600a7d6de7740c8c13f01d1703b32

Observation d2dc4718-1a32-4d2a-952e-3628a94ac764 · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T20:26:13.135594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:e39d4f75969465fe5e2adad6f0e09a3a994be198b8547ab81b271c6c1391df24

Observation cfd54f65-456c-433f-bae3-90776aa463a9 · inbound

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion cites this paper.

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-06-28T20:52:37.897010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T20:44:20.190064Z digest=sha256:aa8c12686881f40e6d74e1f7d22406ef558b248b8bbcc85b909482cc6c517283

Observation 27dcc097-dae4-42d1-b873-027a5d2b0953 · inbound

LaSR: Context-Aware Speech Recognition via Latent Reasoning cites this paper.

LaSR: Context-Aware Speech Recognition via Latent Reasoning Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T19:22:34.997415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T19:12:59.843096Z digest=sha256:0ee469cf5274a59fa922f6ca257d0842a015a170fe1b2d2ea0e9abb89172aa4c

Observation 791d3621-378a-4a24-b29b-8342366a8cc7 · inbound

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects cites this paper.

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-28T17:52:27.089254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T17:44:07.669223Z digest=sha256:94fe0bbcadf8a32c4346d1bc664dd10593844705577efb044e2cd5b3262547bc

Observation d81efb37-1666-4930-9056-de1de2e10669 · inbound

UniVocal: Unified Speech-Singing Code-Switching Synthesis cites this paper.

UniVocal: Unified Speech-Singing Code-Switching Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T00:46:24.621193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T13:17:13.510587Z digest=sha256:bab5749d66bb83d90ba30e9f32beacac63213b0f852e2e5142d408f8d858b9c1

Observation 697c4449-917b-41c8-bb3a-db42a690e70d · inbound

EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement cites this paper.

EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic Entanglement Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-28T12:42:08.977406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T12:34:06.024192Z digest=sha256:4c5ec68adb2cda94aeaa5747d68dea3bd6fb67b38ea0a3e3ab46c8148156c22b

Observation 2ce5ec35-c2ca-49f5-a2e9-a3785716a6e5 · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T05:16:39.811706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:ed418b6a0eb4a96259881bb4159ed82aa3e11f25750f1b26c360ebccc1e8469b

Observation 99a214d0-0985-4774-a150-9dcfd89a5d19 · inbound

Foley-Omni: A Unified Multimodal Generation Model from Task-Level Audio Synthesis to Complete Video Soundtrack Generation cites this paper.

Foley-Omni: A Unified Multimodal Generation Model from Task-Level Audio Synthesis to Complete Video Soundtrack Generation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T04:56:39.434750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T08:38:44.906160Z digest=sha256:50ebc4dcf4814047e31c09c6ff857505737c4af43aed3f3d09e7c2a475a3b23d

Observation 56cca99c-ac62-4935-844d-3307dc3e73b2 · inbound

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding cites this paper.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-06-28T05:21:39.872068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:dadc466a94a524ad41246aa0e1ae6018421d96bdab7937a39db32c34fc5fe895

Observation 63bb06b3-a176-42a3-a002-abe3735fdc0a · inbound

GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech cites this paper.

GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T15:27:04.883488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T23:53:55.385445Z digest=sha256:be0fba5f8d41e76d4bb30022c55285e4c8a9450240aec9c1c1a4388a8ba6e269

Observation 03a510e2-8b82-4fb0-801f-762f3e0dd610 · inbound

HybridCodec: Fast Dual-Stream, Semantically Enhanced Neural Audio Codec cites this paper.

HybridCodec: Fast Dual-Stream, Semantically Enhanced Neural Audio Codec Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-02T15:47:06.435229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T23:27:29.942344Z digest=sha256:ea02c282e10e92c40c031ec378ee1856832a050012c94149eabc346e7ae7675a

Observation de8cf5ce-c91f-405e-b045-5ea72d02aea4 · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:47:19.797481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:1af0075817fb3ab6d0f97d801e8ee9dba6ab414de35b964c3aad3b682e886e33

Observation 5826c81b-871c-4474-8f24-be0c20aeceb9 · inbound

dots.tts Technical Report cites this paper.

dots.tts Technical Report Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T19:57:20.079050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T21:10:25.911203Z digest=sha256:8b94f22aa4a9dd4e1f5e3eb35cbbef1c36a6484650bdb645b18eaf29f776b6ec

Observation 6a2f00b2-9121-43f4-a782-855895f90337 · inbound

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech cites this paper.

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:07:35.860043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T15:27:01.142747Z digest=sha256:5871c86ac3f3e7dcdc5f7911801366318c21267a528825cfd1132cafa867c4c2

Observation 08f970a2-ee8b-4e8f-adbe-114e3cb3ea4e · inbound

MeanVC 2: Robust Low-Latency Streaming Zero-Shot Voice Conversion cites this paper.

MeanVC 2: Robust Low-Latency Streaming Zero-Shot Voice Conversion Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:37:35.401677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T15:11:53.320505Z digest=sha256:ebe8603f54f8841edf592339682fe74dc05ed3719d808b44b34b1e52c17f575e

Observation 5ce2933d-a971-4f6e-84f2-e9b06c8422e7 · inbound

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation cites this paper.

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:27:36.133608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T15:15:10.060770Z digest=sha256:4dea1d623e962f8785bb43f4f9bba6fcec862a51f82c1fa7d4cbde0f561f8117

Observation aff23cce-2c94-4fff-a242-33255dcde729 · inbound

End-to-End Training for Discrete Token LLM based TTS System cites this paper.

End-to-End Training for Discrete Token LLM based TTS System Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:27:34.885277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T15:22:06.893507Z digest=sha256:745e9e20e596d60fdc7dd22be7eb68cb40a2a6b78a3dd48b5b0b69681a96f32c

Observation 71d79e18-eb6a-4310-a9ef-82b96131562b · inbound

M*: A Modular, Extensible, Serving System for Multimodal Models cites this paper.

M*: A Modular, Extensible, Serving System for Multimodal Models Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.693797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T10:01:04.153130Z digest=sha256:7d75ce9645d2f8adfd739459aa8b73b43b5d6350fe2ba4d2dcaed801fdde41b5

Observation d9dd68d5-dbbb-48fa-8d26-52246c014db0 · inbound

Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations cites this paper.

Investigating Human-Model Discrepancies in Speech Quality Assessment via Acoustic and Prosodic Perturbations Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-04T05:29:35.952858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T15:57:06.402107Z digest=sha256:973344a9a1e8fd0713d5755416b596f1329b2d8925085aa472a7708b57a0e07e

Observation 9535b3f4-9a54-4aa9-aef8-edd9b2d215f9 · inbound

Zero-VC: Zero-Lookahead Streaming Voice Conversion via Speaker Anonymization cites this paper.

Zero-VC: Zero-Lookahead Streaming Voice Conversion via Speaker Anonymization Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-04T05:39:39.684444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T15:49:21.097255Z digest=sha256:18d24f2a587c57a1fc800c7e0bd55d551e93e6c17385b9aa1e9bb53631cdc2c6

Observation 5e24eb89-96f4-4e72-a11c-44b7b17e5bc6 · inbound

Transcript-Free Flow-Matching Text-to-Speech via Speech Feature Conditioning cites this paper.

Transcript-Free Flow-Matching Text-to-Speech via Speech Feature Conditioning Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-04T05:39:40.767341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T15:38:13.819676Z digest=sha256:639991093bef50ca60330faa4ca41e26398c9d65fa662112bf978bd65177e5f1

Observation 503ee912-f86a-42e1-811e-e14b1b17ffde · inbound

Imitation Learning for Elder-Facing Speech Synthesis cites this paper.

Imitation Learning for Elder-Facing Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-04T07:29:39.025813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T13:24:48.094668Z digest=sha256:e8e3cec3346e1f582f910fe647aef980bcecfbe21650410a79048509527f5f66

Observation d3913959-4d88-4849-840a-12893bb77b8e · inbound

Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption cites this paper.

Bagpiper-Edit: Zero-Shot Open-Ended Audio Editing via Rich-Caption Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-04T07:39:38.482631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T13:17:40.718000Z digest=sha256:c4b044d11427bd0463ea7af79b61fe8e5dd5f868fcae8256cb376cd0415855e8

Observation 4e8f3d24-7753-47da-9c58-309bf989a0a0 · inbound

ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion cites this paper.

ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:19:44.592983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T11:49:04.308326Z digest=sha256:75a21e4c58808c76cc2458b55d32072de96adba0ffc43cf0a5e1a08d86453c4e

Observation c90182f8-5fa8-48f0-9058-3fbfa6ba711d · inbound

Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis cites this paper.

Bagpiper-TTS: Natural Language Guided Universal Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:19:48.001869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:54:18.325054Z digest=sha256:1eb5f6ddf674015bf86d88fb0f02d40a85ca683faeafab3e7e3d0fc6b07e4e89

Observation a6aa5307-c326-4407-96aa-9bb43418adbb · inbound

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation cites this paper.

AudioCALM: Continuous Autoregressive Language Modeling for Universal Audio Generation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T11:59:50.985665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T07:24:01.244733Z digest=sha256:5458f6d5d84ba8d098572e8720f1d0cd74f9b322f7bbf7a5b40a95c59adc0bb9