Pith. sign in

Paper Citation Record · LEDGER

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information

As of 8 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2505.17426.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17426 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:53:18.993310Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved31
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9d33002b-6b3c-47b2-89c1-3fcdaa12fcef · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:14.619002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:14.619002Z digest=sha256:b66fe4f39345eadf35e5e4bcb3246cc54299ae850d9babbdd807d90460752d78

Observation e4d26acb-124c-476a-b934-3f4e48d06b50 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech represen tations.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information wav2vec 2.0: A framework for self-supervised learning of speech represen tations

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:23.656153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:14.670603Z digest=sha256:10cc053e5184ff74ce1fa7cb025ffe5b131244fbf8911f132002d53ddf5ef000

Observation 966c5132-20c0-440d-b084-c0e5db6586e4 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:14.733869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:14.733869Z digest=sha256:9065e8b2d764fc2db09cb5a0b35dd32d6ecf78f52bb3c4d6fb473613aacbe70f

Observation ec869804-b9ce-409a-9c1f-486251b2b3c6 · outbound

This paper cites High Fidelity Neural Audio Compression.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information High Fidelity Neural Audio Compression

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:14.823243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:14.823243Z digest=sha256:3297c44ac1f0624a7a95b3cb14d2cb0182256d722199c453cde927e0fb53c4a0

Observation 213665fe-a09a-410c-8c8e-d40078715e2e · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Moshi: a speech-text foundation model for real-time dialogue

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:14.907708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:14.907708Z digest=sha256:fca34096338baa10c1d4e3136b747bd5aa1aaf72f4849b61d6c19bf716993c59

Observation cceffa41-b9b1-4038-bf68-672cd9a46730 · outbound

This paper cites IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:14.979998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:14.979998Z digest=sha256:6c5a9722204002848fd69a8224c0c355f547c6a25879544ff62a8f3fec7ab88d

Observation e607820d-e8a4-430f-955c-79c2edb8e86d · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:15.064261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:15.064261Z digest=sha256:8a1ce25806fcf08ab7a28e5cd4a125e95e4803c843f2c4be4b198a81788dfc6e

Observation 0facc9cb-428e-4f1e-8901-17226756ef20 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:15.130965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:15.130965Z digest=sha256:b96120d91cbd7ca631acc3b1fbc0f32a0ae4ec26655dd657a0def34e0b272cbf

Observation a113fd37-95e0-4b9c-8a1b-a3282570a878 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:15.212421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:15.212421Z digest=sha256:2a8fdad45eec915f71e60b5a0183878b2a15123f1583127a6cc942b9fa4e13c6

Observation 7b0deeda-b6f5-4529-8ee7-58b20615f6a1 · outbound

This paper cites The Llama 3 Herd of Models.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:15.302882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:15.302882Z digest=sha256:29ea4e894e68fff40ef236d53e9b6d98844d9b40c6fa22bfb207db73c3dcaf3f

Observation 2347bbc7-4675-4758-900f-fef985939400 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:23.450367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:15.463966Z digest=sha256:c5b0b007d5a9e6c1e09eda88de4e81ee39cc2866f64caec4ca47e2170a95916a

Observation a56d4796-c3f9-436d-91ed-da5e3a3e48f0 · outbound

This paper cites Hubert: Self-supervised spe ech representation learning by masked prediction of hidden units, 2021.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Hubert: Self-supervised spe ech representation learning by masked prediction of hidden units, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:23.285384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:15.590347Z digest=sha256:82a34995d23a1d286605936d09479b25a98c5c625668e6267276ea9246f91900

Observation f9ef53f3-404f-46bb-8008-bc5d8b4d1de1 · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:15.691624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:15.691624Z digest=sha256:90dbe68e7777b5dcf2a023be8e0c4739a6b2f119e8c9b53ba498c58846d6c78b

Observation fcdc4326-4c69-4601-a2ae-5b9241fa4864 · outbound

This paper cites Libriheavy: A 50,000 hours asr corpus with punctuation casing and con- text.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Libriheavy: A 50,000 hours asr corpus with punctuation casing and con- text

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:23.159193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:15.795893Z digest=sha256:5471395be094dd8b09347dd7d1bbd03a0e012d44be3be8e6fdcecb6133ed56a5

Observation a3f97cc6-f5d9-4d2f-8b49-b022605b9663 · outbound

This paper cites Scaling Laws for Neural Language Models.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Scaling Laws for Neural Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:15.934554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:15.934554Z digest=sha256:773b47b5aa0beea9cf627b7ddba41f60d53b8b549b469a20963c81412f4e592d

Observation dbc7af43-37c4-44bd-94e4-40c2e31177c9 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:23.005927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:16.069772Z digest=sha256:a76917507e730d495bd69870c086b519818a0034254ae4294590b75062406642

Observation 15d8824e-45e5-40de-add5-0652964ca1f8 · outbound

This paper cites Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.222193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.222193Z digest=sha256:36110118dcc0076ef364a9672fb2c7c108f69fa4a3ff76eae5802e7ec0bfc175

Observation 1bafbe8f-77ba-41e9-bc48-584777613c6e · outbound

This paper cites Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.316201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.316201Z digest=sha256:f351c063d8e548dfca0371f27683504d032290e9b832baf6d4d4feb2a7d4fc5b

Observation d7a1aebb-b7f9-4d63-b941-2dfb44717064 · outbound

This paper cites Unitok: A unified tokenizer for visual generati on and understanding.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Unitok: A unified tokenizer for visual generati on and understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.366904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.366904Z digest=sha256:9ff83fa8e1dc26eead4bc9b51946e0d16906fc77c6b506eb1f261b284aa182e6

Observation 30bb3cbe-ee07-4b8c-98c1-e66a3eca2135 · outbound

This paper cites WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.449290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.449290Z digest=sha256:2c9dacf97cc9afeda28c7ab27e947d873c28c157c8b22c555710857118264eb6

Observation 10d591a9-3c72-4840-b415-418f7995e5cd · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Autoregressive Speech Synthesis without Vector Quantization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.516851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.516851Z digest=sha256:0fed1006aa2a0af157a82517f487a3b26e84768f6145dfced2918a64e75a96a6

Observation e04f95dd-98d0-4488-a399-b92592437938 · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Finite Scalar Quantization: VQ-VAE Made Simple

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.616864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.616864Z digest=sha256:2e72c4adf408e598f5ea6f37a4c7cbbe82fc44ac4892bcfad5eeb472b99018e7

Observation 53fa332c-d27a-43da-966e-4b834a73cad8 · outbound

This paper cites Scaling Transformers for Low-Bitrate High-Quality Speech Coding.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Scaling Transformers for Low-Bitrate High-Quality Speech Coding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.774762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.774762Z digest=sha256:bc5142de8fe84e114765d64d2f1d64e523c2240094a6252415222b930e9634f1

Observation 72b749f9-c540-4490-8fa4-134eb682d2ac · outbound

This paper cites Loss-sensitive generative adversarial ne tworks on lipschitz densities, 2018.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Loss-sensitive generative adversarial ne tworks on lipschitz densities, 2018

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:22.341747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:16.862660Z digest=sha256:a46b20f495c2542bae218e6fb1c59d4cb1a420dc5d82f13b542f213ff7e7862d

Observation 87751d8b-54bb-406f-aeb7-ba408ac83444 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Robust speech recognition via large-scale weak supervision

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:21.518013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:16.955473Z digest=sha256:94f9c0929b2e2c65d07a228c8151bffb52a565e840b7aa4e8176ad041cdff460

Observation dc727f44-2f28-4626-9890-34064b2961a3 · outbound

This paper cites Language models are unsupervised multitask learners.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Language models are unsupervised multitask learners

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:21.364048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:17.076214Z digest=sha256:107f227f0717d638883885b71b868d1586cbcee686e71a22725e3862b7f3dda7

Observation f120fd47-36e9-42e9-ae14-53b979da2abb · outbound

This paper cites Direct preference optimization: Y our language model is secretly a reward model.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Direct preference optimization: Y our language model is secretly a reward model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:21.223798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:17.171284Z digest=sha256:6d09e8f153742f0190048e687d05e9672a5088bdc1bce910b0c594bda65d2840

Observation 9701ba54-0ecd-4c53-9aae-316cfd4e8577 · outbound

This paper cites Dnsmos p.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Dnsmos p

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:21.106441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:17.261661Z digest=sha256:9681c15eaab1211f2c06cb8745e0db9d0720f60fe26afbd56e5ec5b04df781e2

Observation afef3ee6-190a-4445-9f43-6feb6b7343c1 · outbound

This paper cites Utmos: Utokyo-sarulab system for voicem os challenge 2022, 2022.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Utmos: Utokyo-sarulab system for voicem os challenge 2022, 2022

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:20.984555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:17.331644Z digest=sha256:26688c6e13285cde174c185aadf84db0916fb5d5eb15d7fbe3740c643bbd158c

Observation 66a420ae-6782-43af-96c9-7844c3edf45c · outbound

This paper cites Neural discret e representation learning.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Neural discret e representation learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:20.810423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:17.426791Z digest=sha256:30a4e99d50c37e8e00669bfa5837ff4006d3a9c14bac84101bb5cc9f66776037

Observation f6fd71d8-af69-4bf9-956a-c795fc93f5f9 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.522496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.522496Z digest=sha256:514f1163c0f52877e2a262f74383dcafd8e25c510cc2099ec23219c55a08d2bb

Observation 2a745d27-1788-4f6c-9bd8-47e1b4265e34 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.592637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.592637Z digest=sha256:714947441091887915918d717014c9161b7424bceda1057c3fb625fe00fb3914

Observation 6b94630e-26d1-41bc-9996-33324ce4bddc · outbound

This paper cites Convnext v2: Co-designing and scaling conv nets with masked autoencoders, 2023.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Convnext v2: Co-designing and scaling conv nets with masked autoencoders, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:20.707290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:17.651875Z digest=sha256:a8f411b6ff33856be101adfba74af7c1e4fff862aac16605ea4f04a3e7131add

Observation bdc4cb98-4edf-4f45-80c6-4ec8e460e120 · outbound

This paper cites BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.728250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.728250Z digest=sha256:a3a16bf3bfe8e91850687c205a58fa80cbc02728ad5c6a62c2d9c6720ba131de

Observation 0abc68c5-c879-4e25-ab80-8811ab6e65ca · outbound

This paper cites Qwen2.5 Technical Report.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Qwen2.5 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.785623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.785623Z digest=sha256:a7115658a8b25ea54a8b753a43ba509c13fc71ef6da71d808540a8c8984ff07a

Observation c17515ea-574e-4cd3-9cdd-f3aa64e2267c · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.856062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.856062Z digest=sha256:26d8c799682fc436c3ff8b3449dd3dd0e3fb46972c3c97fa5ebf94c68a520c21

Observation e3ba37b1-24a5-4a5c-a7e0-b44ff18477d7 · outbound

This paper cites Codec does matter: Explo ring the semantic shortcoming of codec for audio language model.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Codec does matter: Explo ring the semantic shortcoming of codec for audio language model

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:20.545166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:17.908432Z digest=sha256:25d22d4dc01afae6a1dbe763c442e5488c4508ad6cb26e478868b561e9991b5e

Observation 7f6a92ec-8d1c-4ad1-af56-c56ba9d1e6b1 · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.983678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.983678Z digest=sha256:9e05cb067d1a9d0b884a57e8e8dd39343fbcdaa67e188d43cf007307abaac518

Observation e4532399-60e4-4eed-a94b-1a509428d199 · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Vector-quantized Image Modeling with Improved VQGAN

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:18.048407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:18.048407Z digest=sha256:9ced15afbb4d7142dd5ed58df2a2a8e5298702b3dbaaaf97b126bb3aaa4a1957

Observation 87603f81-26c6-43e7-9b68-d341b58d5cb1 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Soundstream: An end-to-end neural audio codec

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:20.421251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:18.148272Z digest=sha256:e83b13d72a87cd1c70fb250ef3f4b90220ced1a89f4a7d74f5a095ebae9cbbc3

Observation 1af4cf71-c62d-4614-881a-7eb6c32833fc · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:18.271919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:18.271919Z digest=sha256:43c5f19bccb03be1b869d7a91a88ecaba38e566efa81f9e91f08091e43587af6

Observation dee18e11-8a7f-4bf7-bd7a-55144e5b916e · outbound

This paper cites Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:18.403349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:18.403349Z digest=sha256:67d835460025cda07f2fc907b77c9db1a005ccf136b05175347819caf8524293

Observation bf398295-290a-4cd1-a408-de9d69fe6597 · outbound

This paper cites Autoregressive spe ech synthesis with next-distribution prediction.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Autoregressive spe ech synthesis with next-distribution prediction

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:18.507278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:18.507278Z digest=sha256:b82817ac1c41580bbc4604dc029c8feacc7a696f88705df773ff257466acc6ac

Observation 40441d08-839b-4330-9113-0461d391477a · outbound

This paper cites an unresolved cited work.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:53:20.293072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:18.615503Z digest=sha256:4fc190449f67701d63cfcd08154e8c1e80f0cc40dfc6d6ad7c7cb36e1011dd7e

Observation 409776ea-b68d-4b8a-874b-af7b89ec1c41 · outbound

This paper cites an unresolved cited work.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:53:20.071455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:18.743653Z digest=sha256:11a777213a1477bfba6f1dd8671b26f18c47aac809481e05e57318f67ccd4490

Observation 04b30b9b-3f9c-483b-ab74-e59c35936428 · outbound

This paper cites an unresolved cited work.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:53:19.867514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:18.852069Z digest=sha256:5ba910a8f2994d1f0ed7407e620237672c83dda5bad6c6fd1f726f63e1b13f7a

Observation 173d7a8f-9474-45ea-9412-76d0b909fe51 · outbound

This paper cites 16 Table 13: Multi STFT Discriminator parameter settings of Di stilCodec.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information 16 Table 13: Multi STFT Discriminator parameter settings of Di stilCodec

Reference 47

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:53:19.646311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:53:18.993310Z digest=sha256:be2aa6e7e679cb105bd124d151743475444a31c4a0203bfed7efa218a78611d4

Pith citing papers

No inbound Pith citation observations are available.