Pith. sign in

Paper Citation Record · LEDGER

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations

As of 7 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2507.12197.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12197 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:55:50.720088Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 971e9b2c-bccb-40d3-ac1b-bd6a2a5cb146 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.452658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.452658Z digest=sha256:2d898afe16816627e26544c0e944d0b1b680edf2d452bf5229cd4da5128269e0

Observation 90cb3af2-4b64-47b7-972a-918d54202787 · outbound

This paper cites High Fidelity Neural Audio Compression.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations High Fidelity Neural Audio Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.747483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.747483Z digest=sha256:4338cfc5dee6ae7b8ecc96f139678780ca95a54f604ff3d0d6e458f8cacff697

Observation 70fe884f-f9ee-406b-9b21-b6aaa865037e · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Moshi: a speech-text foundation model for real-time dialogue

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.826784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.826784Z digest=sha256:7f71a2f0f35eae0b401ac420afa4cce096b374d8a664933a0d10ee32a46b85bd

Observation 8ee08fab-3606-49ca-a030-696e0f56c49b · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.922280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.922280Z digest=sha256:bf578c8c8446f8b8c167531d26777ffe636be059354f077675956f9cdaa6f54e

Observation 06c0925f-3d35-4bb7-9e45-55b0626476e6 · outbound

This paper cites Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.981725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.981725Z digest=sha256:307860ab05e725d77979167bff4078ce7c2e01274e0e78125bdf9d7cb9f48577

Observation 5cf50fe2-f983-4810-9806-b75e9749fe36 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.078904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.078904Z digest=sha256:30fa19b3f3dd958a86fbaf46792c971fceed5aaffe54702a3a958211c5b95962

Observation 87337e4f-1806-45e6-bfb4-bede7ddf54e3 · outbound

This paper cites StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.153909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.153909Z digest=sha256:d9be3339d9c934e5f583deb64ebb9dd70224a0e031359549f5326c7587cdf38f

Observation 02ec6650-2b1e-4979-8b8e-003742c21c5d · outbound

This paper cites Scaling Transformers for Low-Bitrate High-Quality Speech Coding.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Scaling Transformers for Low-Bitrate High-Quality Speech Coding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.237226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.237226Z digest=sha256:d4a79d074bf18d9aed680425521d000f0c0ffd95c7b0a6dab718194d62bfed31

Observation 628de9bd-5f27-43bd-ab8d-750b22b24240 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.287364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.287364Z digest=sha256:fd6b5a726ae89a2e5c571baec29ab9b3356d5d38e73975a6ab8ebbd170afd756

Observation 506e6954-1a24-43c5-aaa2-a2c3c992b00e · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.479314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.479314Z digest=sha256:86a8bb09d9c123434a1dc31f1f7cbf24877a1963e48ef82af11b8f6e2944235c

Observation 8df2d612-4795-4727-945d-4dbb76992aac · outbound

This paper cites Multi-band melgan: Faster waveform generation for high-quality text-to-speech.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Multi-band melgan: Faster waveform generation for high-quality text-to-speech

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:55:51.136842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T16:55:50.545372Z digest=sha256:dc63df9eb38d22ab717f19a15c3fa51b70d811aff4b79520e72a251ddf11ffc9

Observation f0330872-1376-4cc9-92cb-e1e56c0a857c · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.665549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.665549Z digest=sha256:41cd504e036550a2cd9ceed3db5b6eb98e4242745cb5ee8a642539705dcc4ea9

Observation 5a717487-dfee-48e8-8cad-7c7e4e17f5a7 · outbound

This paper cites MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.720088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.720088Z digest=sha256:1923987a267bae5f6ffb9cbf470c67755d6ba3b0953845bb97e7026bfc31840f

Observation 45171cba-3a9e-437e-8d8f-363ea2e819ae · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.034113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.034113Z digest=sha256:ae7ac150f8fcd3a0fead9441a40126ed3130ae9368cb9abef95b63228fd052aa

Observation affebd7b-7615-4ba4-9882-6e541db2e17b · outbound

This paper cites doi: 10.1109/jstsp.2022.3188113.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations doi: 10.1109/jstsp.2022.3188113

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.676482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.676482Z digest=sha256:b2d38578374a94eee241e21dbb8b7dc40207fbf45f98e6f65c24a556c934615e

Observation 799f044e-036b-4f20-a99e-2041da5c1500 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.583767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.583767Z digest=sha256:f81bb1daed2ec4f4987c381455c1ddb8bf8f5fac38b04753a672afd697304dcf

Observation edcddf70-53f9-47ca-b41a-b5fb3e02e6dd · outbound

This paper cites Better speech synthesis through scaling.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Better speech synthesis through scaling

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.512702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.512702Z digest=sha256:cad592b0147d3462487058940530dcbb34c00a314759b3493576601015619565

Observation a32f064e-3d61-4988-8f59-e4bee014f2fa · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.368580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.368580Z digest=sha256:a03eaa76e62cc3c9208e7278ccec661b29e91f2ef34812f32e69b194d6d0b525

Pith citing papers

No inbound Pith citation observations are available.