Pith. sign in

Paper Citation Record · LEDGER

Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2306.15687.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.15687 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:21:33.236994Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:47:06.014261Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f456c7f4-541e-47cd-856c-1b8ad3962f2c · inbound

Movie Gen: A Cast of Media Foundation Models cites this paper.

Movie Gen: A Cast of Media Foundation Models Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:25.046954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:3a26eed5b27a1d7426d6c7c665ad25a60b36516c5bfccd7a261054a46239033d

Observation f142280b-7379-477a-bfee-909d5bfdb87c · inbound

Unlocking Speech Instruction Data Potential with Query Rewriting cites this paper.

Unlocking Speech Instruction Data Potential with Query Rewriting Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:21:33.236994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:21:33.236994Z digest=sha256:73c2e6358d7642d98ab0909ecc71d90ffda4a69ed0f43439cc8d09f4e7b60823

Observation ead94f66-2b8d-4b78-928d-1d85dd3b77a5 · inbound

Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages cites this paper.

Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:30.133816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:14:30.133816Z digest=sha256:e8c6ae30a6fa774765b6eb8cca35dc37d655b399998a1785def3e31ce2091c40

Observation a26f963b-c51b-4469-a6b0-76b0961ebf9d · inbound

A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction cites this paper.

A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 119

Resolution
unresolved
no resolver link, observed 2026-08-06T13:54:39.883971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:54:39.883971Z digest=sha256:1704f21f6fc662c8131823499a16cd0e679b8be53975080aee5ecc4cb9e93130

Observation 90acce7b-57d5-49f2-9b22-d93e9238a3c7 · inbound

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training cites this paper.

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T18:03:40.202491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:03:40.202491Z digest=sha256:534614e2f846b573224b0f21424e4c35d50070c4922a4942a33b2b6a616fce85

Observation f9d4e9c5-ff4c-451c-acc1-e5a81b6d86e5 · inbound

F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation cites this paper.

F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:47:06.015877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T23:36:27.369551Z digest=sha256:ce5d4fbd3649134fd63f093060258295160f1929546544df80992556e50eaa95

Observation 02bcab2f-d7b1-462a-b062-50ba2e73d25c · inbound

Optimal Self-Distillation for Rectified Flow via Linear Probing cites this paper.

Optimal Self-Distillation for Rectified Flow via Linear Probing Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T00:41:47.855999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:41:47.855999Z digest=sha256:92988fc4660dd0b9251a4cfbfaa697fb29f9fa6594e6d4835a463a0cb7af3391

Observation b386eac8-8089-4c0d-89f8-1bac72d4b507 · inbound

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System cites this paper.

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:41.413951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:41.413951Z digest=sha256:43a9a03064e95c91676513a9ac692a05846e3cc6e3b9e369e9311d901e340162