Pith. sign in

Paper Citation Record · LEDGER

BigVGAN: A Universal Neural Vocoder with Large-Scale Training

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2206.04658.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2206.04658 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:10:56.252891Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T05:39:39.351499Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d23726c2-113b-44f4-aa51-28e4166d5cf8 · inbound

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models cites this paper.

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:26:37.374042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T12:26:37.300599Z digest=sha256:f816dcb2afc5af45a2713715b1dde45c45f6552943e53c251dd8d1c078820d48

Observation 2739dbb6-bbdd-4204-ada2-6c2aa237b80e · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.419455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:19955d1dda9e508c1e7610e063d18b5eee319ad159bdfcb23735a90e56145e82

Observation d1f8a6af-4077-4cfd-922b-e4af6a01f3a2 · inbound

Movie Gen: A Cast of Media Foundation Models cites this paper.

Movie Gen: A Cast of Media Foundation Models BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:25.090823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:b85dcefc734218919511c1cd9fbf979320ddb6fa9b8f2d5c9db6750157f0613b

Observation 7d63e650-8a76-42e2-8d3b-161003e822e3 · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.203852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:ff4b98c4c4a7a54c5c8c588e70c446991ec069e9d71872ba34f45ca19993f9ea

Observation 0b615ebf-df7c-41fa-8796-d91eb89e070b · inbound

SwitchCodec: A High-Fidelity Nerual Audio Codec With Sparse Quantization cites this paper.

SwitchCodec: A High-Fidelity Nerual Audio Codec With Sparse Quantization BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:12:18.182021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T13:10:14.839742Z digest=sha256:8c605daf25a66b5e16d1194802c58c445802d3dc40307d86acec114d30c83b5f

Observation 604528b6-756c-4732-9d83-2299a1a9d949 · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:51.180223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:29699cfce96b229ce8020cd2009c39a03c40c2147046dcc6c363ccaa7980a12a

Observation 1df455ea-20ce-4795-9461-174ffa8818ab · inbound

NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference cites this paper.

NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T23:10:56.252891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:10:56.252891Z digest=sha256:47d3932e74717a5837a4f120eddf384ff569e6d8ea84a034cdffdefbbb180454

Observation 1c4df732-860b-4121-aaf6-1bbac2558344 · inbound

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation cites this paper.

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:03:15.250457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T01:03:15.183360Z digest=sha256:051efeacb83e27a6d483e9bec3e64540e1772bfd45cc774e33bb2d9bc3be5bb0

Observation 373248fe-be09-4514-ab30-ec207ad538f0 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:33.696265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:33.696265Z digest=sha256:38fad697c6a0f4c6ba0e074425308b8fbad0024b67aa826ac8ec2e327a7f13fc

Observation 0a9243e4-5b1f-4c08-8647-a0501dc8fe38 · inbound

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model cites this paper.

OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T00:11:13.570717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:11:13.570717Z digest=sha256:773f6a9f2aeb9571f64b97ae8a457a4ace8523ef22adfea40b587f297ef052d9

Observation e6e78f76-5b58-4f8a-ba39-c53405f146b5 · inbound

SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns cites this paper.

SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-14T22:43:13.611383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:43:13.611383Z digest=sha256:3373f71614446a277583d4047335288c1b63d5d773a59794fa9435d84c8b99fb

Observation 226673d6-99e2-4c78-aaa2-02f64a4e0c1d · inbound

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training cites this paper.

Making Separation-First Multi-Stream Audio Watermarking Feasible via Joint Training BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T18:03:42.596651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:03:42.596651Z digest=sha256:6fd07ea39fa3775f6cc85d7b2b68e9d3a41fc4fa1f86e3b13801a67c7af1eef3

Observation 20153b22-c0d0-49e9-8ffb-4c85d2fbee4b · inbound

Latent Fourier Transform cites this paper.

Latent Fourier Transform BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:21:07.071659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T03:45:07.892234Z digest=sha256:3061365ad84412afee3ffb1a5aec65c9300f59117a6a80d78b6dafa18b52d816

Observation 9992c7a6-1cbf-4232-bcf4-a04420ee13d6 · inbound

Modeling Music as a Time-Frequency Image: A 2D Tokenizer for Music Generation cites this paper.

Modeling Music as a Time-Frequency Image: A 2D Tokenizer for Music Generation BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:47:43.077570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T18:47:32.003545Z digest=sha256:561557b69e777b935f0367310d96a2c099f3d2fabd50009ada7a0486718bd763

Observation 67134491-ea93-4cb6-a203-9f53492153b3 · inbound

A Survey of Advancing Audio Super-Resolution and Bandwidth Extension from Discriminative to Generative Models cites this paper.

A Survey of Advancing Audio Super-Resolution and Bandwidth Extension from Discriminative to Generative Models BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:27:49.816692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T20:26:59.049472Z digest=sha256:f32946603560229bc2c02f32c866c2616d5afc40c774420fa7c2132d96f2791e

Observation 636c0cda-1aa7-491a-803c-a712e014276f · inbound

Taming Audio VAEs via Target-KL Regularization cites this paper.

Taming Audio VAEs via Target-KL Regularization BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:53:23.209087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T14:53:15.718359Z digest=sha256:317e27287cdac8a96cd7ac3739773d9037bd6bf80ff142b8c49bf7c4c8af616b

Observation 3662da70-f9dc-47ec-9904-b88ea400f5ca · inbound

WavFlow: Audio Generation in Waveform Space cites this paper.

WavFlow: Audio Generation in Waveform Space BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:38:09.595548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T07:33:35.243337Z digest=sha256:adfb74e0853cd2cf3a658f50fd676e6ac5abf1680af3ba4f4e2829aacafd7aa3

Observation bf7b9118-f04a-47bc-83e3-a35cb284d05d · inbound

Can We Hear from Events? Generating Speech from Event Camera cites this paper.

Can We Hear from Events? Generating Speech from Event Camera BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:43:30.664608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T14:40:29.819278Z digest=sha256:1d3475827c092bcc69ce2d3b7af861783a91338515993c848d369634f78f8dd1

Observation 0309b8d9-1088-4116-bfbc-45332e306b20 · inbound

Audio Deepfake Detection with Half-Truth Localisation Using Cross-Attentive Feature Fusion cites this paper.

Audio Deepfake Detection with Half-Truth Localisation Using Cross-Attentive Feature Fusion BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 2

Resolution
malformed identifier
arxiv_id, observed 2026-06-29T06:03:08.305425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T06:02:59.488895Z digest=sha256:1f6ec5b23e7a8411c836391cddddabc08bad53cc3f19d3310a229b08c4253f09

Observation f58b15e8-af53-49b9-9393-c734f69dbe9a · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.735273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:ffb11c8198eb9f706bdb4bcecf9420db05a7141fcac81bb591eded230ae8fc43

Observation 6eca4175-8429-414d-9b49-d8ba54451670 · inbound

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers cites this paper.

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:39:39.354621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T15:51:48.672086Z digest=sha256:73acdead162e53b1da41fb37ae7ca1315bf0ac2c3ca3380aa8d10f731bceee53

Observation a817e62f-41a3-4069-9317-bef083050635 · inbound

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers cites this paper.

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:39:03.049720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T23:36:39.368933Z digest=sha256:1a308fdcf4491183c1abef86c239836292b54a96afa1f8a00d94958b4051f6a9

Observation 6330ad10-3a6c-4bb5-a416-99b113edd190 · inbound

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers cites this paper.

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-15T10:41:34.347336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:41:34.347336Z digest=sha256:13fb791fd2714525505ea1b5dbb0d550c10929a54c8d57114d46d6fe3ff4bec5

Observation 84d3053b-a9ec-42ee-b070-7879ea0cfc7b · inbound

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling cites this paper.

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:45:46.188575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T04:05:27.343684Z digest=sha256:d7b20e5ae5d8b66a52698ed8c9f4bdd7416615e87c4f0f28f0beafc302987d6d

Observation aa50874b-2739-4ff1-a3f3-aeaf04927119 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 118

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:55:42.258516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:d129ff371af292292b90c968b680a41591351ce9c230b21fe1726fe5786ec75f

Observation e71e3b9a-8de6-4cc1-99c5-688362adba5b · inbound

What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection cites this paper.

What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-14T14:47:58.592181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:47:58.592181Z digest=sha256:1f3bcc86791453a9f6a687be1d10554b3e2ada621d86ad229ffb4bfbd3d17e3c

Observation e7d11ceb-6de3-46b4-b475-c12e180b5c19 · inbound

A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors cites this paper.

A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-01T22:37:47.065781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:37:47.065781Z digest=sha256:030318fa39c9b28f682e095dd7597c246ebb3f1a12721f650e0a407128e033bc

Observation a2325b15-5f9a-46d0-8265-9b3319349523 · inbound

A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors cites this paper.

A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-04T01:43:21.961191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:43:21.961191Z digest=sha256:9fc7970cf0b64e93922314a65537d9c6018312f0e65b035f9f1f3f4a800ed79a

Observation e5d0f3a8-6f14-4794-95e9-61cf4c3f3ecd · inbound

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs cites this paper.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.851867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.851867Z digest=sha256:3a24bb65053e1aa8d5e9a114694010889c56595fcab113ebfe26595c5b918a5f