Pith. sign in

Paper Citation Record · LEDGER

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure

As of 22 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2608.08286.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08286 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:15:23.055549Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact6
  • verified fuzzy1
  • unresolved15
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b750918b-8b95-4e55-9042-cf6776c0b7fb · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Moshi: a speech-text foundation model for real-time dialogue

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:22.966458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.966458Z digest=sha256:170123b0ed084231364088c88006584f726fb80584c614f5162d265d8c5cd1fc

Observation 2cdbb5ac-acb0-4bf6-b5e8-1379f62bb23d · outbound

This paper cites an unresolved cited work.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Unresolved cited work

Reference 6

Resolution
verified exact
raw_fallback, observed 2026-08-12T00:15:23.716930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:15:22.969847Z digest=sha256:5d8ef6c4a33c196f2888160604117066433b641f43543d2b7d9b2ac384c29f09

Observation 4cafe02c-b60c-4800-9652-81b9419f94f3 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:22.972684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.972684Z digest=sha256:72243230bcb1d60ee171acddd7492dc6bbe81cf1749cf097b87db319b4610e16

Observation 85648047-3cb3-4e35-b1c2-735661c22afe · outbound

This paper cites 16 YiweiGuo,ZhihanLi,ChenpengDu,HankunWang,XieChen, and Kai Yu.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure 16 YiweiGuo,ZhihanLi,ChenpengDu,HankunWang,XieChen, and Kai Yu

Reference 8

Resolution
verified exact
doi, observed 2026-08-12T00:15:23.116937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:15:22.976967Z digest=sha256:2409ee80a666e5a38839e35733e8c9e2382482f1d419d5679bc3329064375b39

Observation c7876959-d0f1-4515-8286-e7a222a7e710 · outbound

This paper cites NadavHar-Tuv,OrTal,andYossiAdi.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure NadavHar-Tuv,OrTal,andYossiAdi

Reference 9

Resolution
verified exact
doi, observed 2026-08-12T00:15:23.106902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:15:22.980061Z digest=sha256:bbc959debcb262af8de5c4f4c978a4b3c478213a0bd28523d7fafd65555d1ab8

Observation 78c4e1da-8186-4199-81db-0b9f26c668a6 · outbound

This paper cites Harry Julian, Rachel Beeson, Lohith Konathala, Johanna Ulin, and Jiameng Gao.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Harry Julian, Rachel Beeson, Lohith Konathala, Johanna Ulin, and Jiameng Gao

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:22.986532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.986532Z digest=sha256:6c2e31e5cea02a441009f0b5dd2ed05b3454c9e8e93cb200a957d5c9f0781f36

Observation e7566782-808e-4cb6-ab93-18919a4410d6 · outbound

This paper cites Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:22.989838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.989838Z digest=sha256:c1754b27edf7a6f13f2d848cdc135a69b991b6f2d2d112a10787a0e2b02b54ae

Observation 3791d20b-f13b-4c85-addc-9828bfc49cd0 · outbound

This paper cites Montreal forced aligner: Trainable text-speech alignment using Kaldi.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Montreal forced aligner: Trainable text-speech alignment using Kaldi

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:23.772910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:15:22.994783Z digest=sha256:eb5a0305b5452e3cd0861b5d98fa9e717c023d0208092d9dcf4b93029e9323ef

Observation e61731eb-39d7-4959-a2b6-697f3f267a64 · outbound

This paper cites Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur

Reference 15

Resolution
verified exact
doi, observed 2026-08-12T00:15:23.082455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:15:23.013594Z digest=sha256:bb006b018d23ce261647f0dc616baf6647a0802c63dc94330889f86a1b58c5c3

Observation 78820abc-3b26-4e8f-b799-c9ffb95c07a7 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Robust Speech Recognition via Large-Scale Weak Supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:23.020983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:23.020983Z digest=sha256:7eb9b441f4def0ac119de479b7bfc76d1ae98dadb5dfd5ec129811fb83f80139

Observation 1e4b5ed5-0c19-49f4-8d7f-5b9810faf0ec · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:23.024338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:23.024338Z digest=sha256:fb0a7142f03e91bd6da6f12d5a90e19f0b7df67760b55465574d13e444d42886

Observation 0e13cfd5-72e7-48c3-8426-d5203d46a295 · outbound

This paper cites MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:23.032346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:23.032346Z digest=sha256:cbec64b82bf1be6beb97a2a9f1af5b7fe0929df300745543ade44a4d3d8089d4

Observation 8f600229-0b91-43ba-9765-7dbfcefa954f · outbound

This paper cites Content is What Remains: Invariant Speech Tokenization from Parallel Utterances.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:15:23.478725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:15:23.039683Z digest=sha256:975820adccaf6a01b04c329a1671ad58b1acad29ba695a5d70cd13c87f8c7568

Observation 19fe2a0b-5d0b-430a-9885-8f3e678956f9 · outbound

This paper cites Yuancheng Wang, Haoyue Zhan, Liwei Liu, Ruihong Zeng, Haotian Guo, Jiachen Zheng, Qiang Zhang, Xueyao Zhang, Shunsi Zhang, and Zhizheng Wu.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Yuancheng Wang, Haoyue Zhan, Liwei Liu, Ruihong Zeng, Haotian Guo, Jiachen Zheng, Qiang Zhang, Xueyao Zhang, Shunsi Zhang, and Zhizheng Wu

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:23.042977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:23.042977Z digest=sha256:e9102eb3b26e989f16ddadb4b470b10477e97931dc437d339f1be39471cdfaed

Observation d910d2a4-a97d-46fd-b6af-db581bfa13d3 · outbound

This paper cites Detai Xin, Xu Tan, Shinnosuke Takamichi, and Hiroshi Saruwatari.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Detai Xin, Xu Tan, Shinnosuke Takamichi, and Hiroshi Saruwatari

Reference 23

Resolution
malformed identifier
no resolver link, observed 2026-08-12T00:15:23.046190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:23.046190Z digest=sha256:4cadfb55e7cdfcbab65cc14f9241063958c1ebb1710cdda2b24a9e8f334ce528

Observation de98056c-99a6-4a4a-ae4c-b4ef6728cb0d · outbound

This paper cites An Yang et al.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure An Yang et al

Reference 24

Resolution
verified exact
doi, observed 2026-08-12T00:15:23.320953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:15:23.049313Z digest=sha256:2f6b350046ac49f3648c02402165fb9a17c2bc39796dc98fd6450531126e89b4

Observation 51606635-d0c9-4acd-8ae2-23c75d0c566c · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:23.052303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:23.052303Z digest=sha256:68f2889219629f1e152536d4abf803764de73885236661f0edb74d87ee92a13e

Observation 1ddde720-ecad-422b-8964-00b9c9ef1c70 · outbound

This paper cites Xin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou, and Xipeng Qiu.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Xin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou, and Xipeng Qiu

Reference 26

Resolution
malformed identifier
no resolver link, observed 2026-08-12T00:15:23.055549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:23.055549Z digest=sha256:4b43ab53d85caee9d91d4ba87f11570c3db811209ffbff505e47e8413228982c

Observation 3c3ffca4-fce9-462f-916b-a4fb3d0967ea · outbound

This paper cites an unresolved cited work.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Unresolved cited work

Reference 2015

Resolution
verified exact
raw_fallback, observed 2026-08-12T00:15:23.598627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-12T00:15:23.017011Z digest=sha256:fc57b8c82af36424d6ebcdf11eddebc88079c5064da658cb89b1c926671ea8f1

Observation e3db7876-5287-4fe0-ac6d-b954b0f0a347 · outbound

This paper cites Pooneh Mousavi, Jarod Duret, Salah Zaiem, Luca Della Libera, Artem Ploujnikov, Cem Subakan, and Mirco Ravanelli.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Pooneh Mousavi, Jarod Duret, Salah Zaiem, Luca Della Libera, Artem Ploujnikov, Cem Subakan, and Mirco Ravanelli

Reference 2017

Resolution
malformed identifier
no resolver link, observed 2026-08-12T00:15:22.998055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.998055Z digest=sha256:2aa0e6464661e7092b7e451d4de6d2fb8f105d6f37517062b45c301aa978605d

Observation f6f5a47b-d744-4265-a5a9-2e296c29b069 · outbound

This paper cites Qwen3-TTS Technical Report.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Qwen3-TTS Technical Report

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:22.983028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.983028Z digest=sha256:d1d06b9766131218636ac57a34319806e5c2d52e8cf346856534c5f461c47a15

Observation b45f9ab1-6ed2-4ca9-bedf-ac2352ae781a · outbound

This paper cites High Fidelity Neural Audio Compression.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure High Fidelity Neural Audio Compression

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:22.963072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.963072Z digest=sha256:9c3bdd414d728abd9aac18ffc2b2a11c346607bfa894e126382db47ed95202de

Observation cb07ca61-e82c-45c3-a0a8-9ccbe848488a · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:23.027676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:23.027676Z digest=sha256:45def068bf3034d9930517fc1e570d8b5cf12f5fe88d8cf9442fd97a690356bd

Observation 7e2eb48c-4064-4630-a7c3-89ab44b59d4b · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:22.951195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.951195Z digest=sha256:9ee1dc020191bb04e0aa091d23d3b87083c382b7a5e4d09c69c02c92f4db324b

Observation f7c4e0b9-93a1-44fb-bb2a-67184bdb22dc · outbound

This paper cites Yushen Chen, Kai Hu, Long Zhou, Shulin Feng, Xusheng Yang, Hangting Chen, and Xie Chen.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Yushen Chen, Kai Hu, Long Zhou, Shulin Feng, Xusheng Yang, Hangting Chen, and Xie Chen

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:22.955426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.955426Z digest=sha256:83c1e494f80ea601de8c85acfd0346502787541d24e9008438ae5cb25674e9a0

Observation ff2a4f74-3c11-4560-a2aa-5a0385ce0fb6 · outbound

This paper cites On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:22.959556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.959556Z digest=sha256:00014502f0882ffb6d3962481f16c8b53614b030528e334bc19d03ff0ddccf95

Pith citing papers

No inbound Pith citation observations are available.