Pith. sign in

Paper Citation Record · LEDGER

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations

As of 7 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2507.12197.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12197 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:55:50.720088Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 971e9b2c-bccb-40d3-ac1b-bd6a2a5cb146 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.452658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.452658Z digest=sha256:f29d30226e48e4ab74ba5a4e4d02f32d7fbd628218a7a1e38e2230970e6ee165

Observation 90cb3af2-4b64-47b7-972a-918d54202787 · outbound

This paper cites High Fidelity Neural Audio Compression.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations High Fidelity Neural Audio Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.747483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.747483Z digest=sha256:84de4d376459dfa041d9dc34fc3cf86891f9754fbe47116797308bca9ae18c35

Observation 70fe884f-f9ee-406b-9b21-b6aaa865037e · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Moshi: a speech-text foundation model for real-time dialogue

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.826784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.826784Z digest=sha256:6bc82786136b04028c738a08bfd37d9ce4196a0504b27a5c2054e4ef8929e790

Observation 8ee08fab-3606-49ca-a030-696e0f56c49b · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.922280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.922280Z digest=sha256:75fc774f9af306380fea24e99dae84f1dfee6072702a289032722a7926dd722b

Observation 06c0925f-3d35-4bb7-9e45-55b0626476e6 · outbound

This paper cites Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.981725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.981725Z digest=sha256:f414cce453ede6a37224c11b452cda7b93c507ee008c327d97f146bf58a5d5e6

Observation 5cf50fe2-f983-4810-9806-b75e9749fe36 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.078904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.078904Z digest=sha256:1a60db3f16762a77d247132f3568fb4337526ad8d0139fe4854043d0ebbe9c2f

Observation 87337e4f-1806-45e6-bfb4-bede7ddf54e3 · outbound

This paper cites StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.153909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.153909Z digest=sha256:62d42a154ce36d532a2cfa9da3380445923337be5fd5f3b0302049e853315e7c

Observation 02ec6650-2b1e-4979-8b8e-003742c21c5d · outbound

This paper cites Scaling Transformers for Low-Bitrate High-Quality Speech Coding.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Scaling Transformers for Low-Bitrate High-Quality Speech Coding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.237226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.237226Z digest=sha256:3f19b56603848104469727be4ebed108fe1f05173e3cf2189da611a29b44635d

Observation 628de9bd-5f27-43bd-ab8d-750b22b24240 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.287364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.287364Z digest=sha256:e838e52493574ba612e1861678b6e02772e644ec395208b5ab953f282ca8a81e

Observation 506e6954-1a24-43c5-aaa2-a2c3c992b00e · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.479314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.479314Z digest=sha256:e7b8593e07ee43f4a579bebe7ad5f1c6df4c131dce61375777336ee83b2559e5

Observation 8df2d612-4795-4727-945d-4dbb76992aac · outbound

This paper cites Multi-band melgan: Faster waveform generation for high-quality text-to-speech.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Multi-band melgan: Faster waveform generation for high-quality text-to-speech

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:55:51.136842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:55:50.545372Z digest=sha256:4a1a17c76ad30302e3e7758a5a402ef0ba3e28f5e64be66da9d0273a93b59223

Observation f0330872-1376-4cc9-92cb-e1e56c0a857c · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.665549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.665549Z digest=sha256:3264ad23ad2a9443816fba9a473592c57051e753585a389d45299c58561ba6c0

Observation 5a717487-dfee-48e8-8cad-7c7e4e17f5a7 · outbound

This paper cites MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.720088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.720088Z digest=sha256:5422863028f58545c3c3be63a540ab4c7fb84f2a59288616c4fcc943202de5e2

Observation 45171cba-3a9e-437e-8d8f-363ea2e819ae · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.034113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.034113Z digest=sha256:7083296fba9023ae097920700d4685dc1f93483232dec021be838b620709dfd4

Observation affebd7b-7615-4ba4-9882-6e541db2e17b · outbound

This paper cites doi: 10.1109/jstsp.2022.3188113.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations doi: 10.1109/jstsp.2022.3188113

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.676482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.676482Z digest=sha256:919e3ba486a29beffbc0ee3ecd466887395c397c7989c90cc94677a24228b41a

Observation 799f044e-036b-4f20-a99e-2041da5c1500 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.583767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.583767Z digest=sha256:2ffe4bdf32d97c6f018ded5ed9f18e4f078cf7dacb51af131826d1d5531572d3

Observation edcddf70-53f9-47ca-b41a-b5fb3e02e6dd · outbound

This paper cites Better speech synthesis through scaling.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Better speech synthesis through scaling

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.512702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.512702Z digest=sha256:210bb5e3431b037a111293a632a2eca066a40ee805353056ef3610362079e336

Observation a32f064e-3d61-4988-8f59-e4bee014f2fa · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.368580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.368580Z digest=sha256:9225e4aabcf5163ac4ca5eb05609fe5ba09993ca3f1866c0e379c47e17b3fb47

Pith citing papers

No inbound Pith citation observations are available.