Pith. sign in

Paper Citation Record · LEDGER

VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 3 inbound Pith citation observations for arXiv:2401.14321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.14321 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 3 of 3 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T06:06:41.348728Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T06:06:41.477473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0e2b1200-720a-4f80-9466-db9ffa8ffb45 · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.479977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:ea46274d522fbec2875093c04ecb76c58973fc5721e4e0a908b49e90efd77fd7

Observation 3498705f-a958-4917-b168-e003ba4c70ef · inbound

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models cites this paper.

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:19:09.541520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T06:19:09.507440Z digest=sha256:748fafff7f90a213d8d4f9be996e6fceac450c09018dfcb6f952850cb592e497

Observation f3c4b565-7901-46b2-9a28-bc5704648bbd · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:25.479961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:13ac364bc7d2878061a494c25c88311f2578651169fc947762f9272065ffeb2c