Pith. sign in

Paper Citation Record · LEDGER

Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2306.15687.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.15687 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:43:40.731660Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:47:06.014261Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f456c7f4-541e-47cd-856c-1b8ad3962f2c · inbound

Movie Gen: A Cast of Media Foundation Models cites this paper.

Movie Gen: A Cast of Media Foundation Models Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:25.046954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:9823a33491db9449b04a651e0e7584434911ba07f4a08d441d3fc8e1ed2d2169

Observation bd427a1c-693f-42af-b43f-14463a3e3cd5 · inbound

Speech Watermarking with Discrete Intermediate Representations cites this paper.

Speech Watermarking with Discrete Intermediate Representations Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T12:44:57.241106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:44:57.241106Z digest=sha256:23e73278367c62efd60680afda4fe4c89f0f79361410775d15290a55a9713c88

Observation b360166c-1f97-4adf-93ca-1ce8e1519132 · inbound

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization cites this paper.

TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.590797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.590797Z digest=sha256:b62018cabf21994ec8a62071d4c859b5e62ed5ddf5eafd43cccef4ab2eb2dd0a

Observation 996e2d5b-fdd1-48d5-88b8-1a4ee0220a00 · inbound

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis cites this paper.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.768464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.768464Z digest=sha256:558f164a76eb11e1901b100356d5256d0aeb8d7d25581cd073ca64f4ddd813eb

Observation 352967e5-f9f1-4b2a-8471-9f1a25b21d67 · inbound

OmniAudio: Generating Spatial Audio from 360-Degree Video cites this paper.

OmniAudio: Generating Spatial Audio from 360-Degree Video Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T11:43:40.731660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:43:40.731660Z digest=sha256:c4b3ed23d53564b03139028654223d36e0f1ccb9177dc37ade0ee2260d4a3302

Observation 8c436648-c066-4e7d-bf43-27b29dadef12 · inbound

Improving Trajectory Stitching with Flow Models cites this paper.

Improving Trajectory Stitching with Flow Models Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T22:12:32.845011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:12:32.845011Z digest=sha256:1742d7461cd9a45e667958ad9c2910af7994dd2f28f3d3d8211cb3a2816c25c5

Observation 228b822e-6e8d-4a88-ad92-ce19c9c4def3 · inbound

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion cites this paper.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.464868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.464868Z digest=sha256:789af2d53390d71bbd80015076ebc14c33974ce13006f4f720484afb9358f3c8

Observation c44eb73f-b23e-479f-822f-8d99bf5d7fe1 · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:51.036153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:51.036153Z digest=sha256:b8185ab98fbfff6c5461e6f3f756d362657d6560da41e9896bb693067f419088

Observation f142280b-7379-477a-bfee-909d5bfdb87c · inbound

Unlocking Speech Instruction Data Potential with Query Rewriting cites this paper.

Unlocking Speech Instruction Data Potential with Query Rewriting Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:21:33.236994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:21:33.236994Z digest=sha256:84124a818c1bb80ac2a4ac8dd853229a2a9d88dd1fcdf35cb4dd7913d29d3d3a

Observation ead94f66-2b8d-4b78-928d-1d85dd3b77a5 · inbound

Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages cites this paper.

Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:30.133816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:14:30.133816Z digest=sha256:92a2ecfb7d180215b5b058abd1044236d7cb27e64e43cbcbe5dbb172cf237e5c

Observation a26f963b-c51b-4469-a6b0-76b0961ebf9d · inbound

Generative Model Unlearning: A Survey through Target Events, Unlearning Operators, and Evaluation Protocols cites this paper.

Generative Model Unlearning: A Survey through Target Events, Unlearning Operators, and Evaluation Protocols Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 119

Resolution
unresolved
no resolver link, observed 2026-08-06T13:54:39.883971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:54:39.883971Z digest=sha256:4f5c2e0fa81fe1894a50fc8b2f643687d2ff262a925f0c4282e6c3b69ad38295

Observation 90acce7b-57d5-49f2-9b22-d93e9238a3c7 · inbound

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training cites this paper.

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T18:03:40.202491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:03:40.202491Z digest=sha256:e6d762f32e745fbe77f477e2de313748eba6fdae7d1cef293fb04be1aebe12cd

Observation f9d4e9c5-ff4c-451c-acc1-e5a81b6d86e5 · inbound

F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation cites this paper.

F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:47:06.015877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T23:36:27.369551Z digest=sha256:c883207928890d47649b3227924401ae9875bc26b4c234bf9f522c5e8000d4a7

Observation 02bcab2f-d7b1-462a-b062-50ba2e73d25c · inbound

Optimal Self-Distillation for Rectified Flow via Linear Probing cites this paper.

Optimal Self-Distillation for Rectified Flow via Linear Probing Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T00:41:47.855999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:41:47.855999Z digest=sha256:30622a19b47439bd9fab8ebb4335da7eddc46e160d02d20db3d3659c02ed3f74

Observation b386eac8-8089-4c0d-89f8-1bac72d4b507 · inbound

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System cites this paper.

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:41.413951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:41.413951Z digest=sha256:7d6cd38e344a65d34ead8d32f7ed86d9a6b9aa12bf4dee96acbf775b7ad9aecc

Observation bbde2eef-b8a9-41e2-9698-44f2f76c5e36 · inbound

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies cites this paper.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:48.017504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:48.017504Z digest=sha256:8831266fc7b655062b1508317ed523788cebb0f0052e1098bcad46b88afd28dc

Observation 0556a7c5-451b-4c06-94bb-3e5db3357557 · inbound

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder cites this paper.

Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:36:21.682711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:36:21.682711Z digest=sha256:c41dfbd67e8364c2e519249993c1615557defe7b2acd129c418c4ec23e6ac53f