Pith. sign in

Paper Citation Record · LEDGER

Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2504.10352.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.10352 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T10:50:49.044638Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T10:37:01.581381Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8e7a93fa-78b1-47ae-93d5-681901ac99a2 · inbound

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis cites this paper.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:49.044638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:49.044638Z digest=sha256:2e5aee05fd83d39967a459682b21011b3071c7cc4f0d878ec09612165b7fdc09

Observation d87052d6-7e3b-4ec3-a499-a0fc5826aa78 · inbound

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching cites this paper.

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:32:03.665690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T04:29:41.285194Z digest=sha256:4ea2ed6c31bf47525fc25ebba9e9885e221207189ff9ba4d7dd573b0c0ad1c83

Observation b2a6f30a-9185-4d39-bc60-d608c1c65b1c · inbound

Next Tokens Denoising for Speech Synthesis cites this paper.

Next Tokens Denoising for Speech Synthesis Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:27.528989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:27.528989Z digest=sha256:c0413dcb6a6546d678ead339d8be7a287a143e869d0254dcad64532bfb3810d2

Observation 5695c82b-c8ac-40e6-96b7-2887b6416d3b · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis

Reference 233

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:55:42.136075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:47cf0df90c373aee8084ab3d3a1778241735aebbc65edce80e87f4168d59d53d

Observation 16fca63d-8010-4856-a07f-35e628cb6480 · inbound

Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment cites this paper.

Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment Pseudo-Autoregressive Neural Codec Language Models for Efficient Zero-Shot Text-to-Speech Synthesis

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-10T10:37:01.582743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-10T10:31:23.170749Z digest=sha256:c303e77f3f754d66d854390131900d51b55a6a79397453c00f5787d9a90d7f96