Pith. sign in

Paper Citation Record · LEDGER

Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2306.00814.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.00814 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:51:35.862909Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

11
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a45dbe55-9889-4076-9708-b7e7239e9345 · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.521748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:64f4451b19a2dcab23678136bca3f0e3ac3c0be2413e4493b2f0e2f17854a6e9

Observation 16507183-20b4-4104-8119-91ffd070184e · inbound

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs cites this paper.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:35.862909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:35.862909Z digest=sha256:6f69aa98bf403f7e30a153e66dc63ab74152f50ee396011a648bba04ed44d2d4

Observation 18874d1a-a21c-4683-8636-141d30755463 · inbound

Traceable TTS: Toward Watermark-Free TTS with Strong Traceability cites this paper.

Traceable TTS: Toward Watermark-Free TTS with Strong Traceability Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:06:18.832332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:06:18.832332Z digest=sha256:3a38597e268bf6532ec77cc4715cfaec6dad8f8a8074b6736b2d289c6695709e

Observation 5b2c4a59-bc6d-40f7-8ffb-889dc66184bc · inbound

Autoregressive Speech Enhancement via Acoustic Tokens cites this paper.

Autoregressive Speech Enhancement via Acoustic Tokens Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:24.295671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:24.295671Z digest=sha256:7adddaa36e76c17de78b2fe74527d1d18562ac8675bfc399cc24d1d365e795bc

Observation 676782b7-0ac9-4013-aa6c-2f897b0f37d4 · inbound

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis cites this paper.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.435245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.435245Z digest=sha256:c15fa88931bdce6e588606c282cda33fa45e42f61ee70fc485acfdd1dd469aa4

Observation 68a7a26e-2885-4c50-9e47-7653efdfdfa9 · inbound

SenSE: Semantic-Aware High-Fidelity Universal Speech Enhancement cites this paper.

SenSE: Semantic-Aware High-Fidelity Universal Speech Enhancement Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:36:22.339938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T12:35:19.023045Z digest=sha256:bd2c532b4183b0ce6c9012ccaff1a1d2f3ea5b3becc4627d95384e1949cdbcfe

Observation 62936d30-161f-4588-b923-eb3c1dbe8c65 · inbound

Two-Dimensional Quantization for Geometry-Aware Audio Coding cites this paper.

Two-Dimensional Quantization for Geometry-Aware Audio Coding Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:20:29.314335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T18:16:51.486807Z digest=sha256:fc4b407ef8f923a3a6e1972d749662fcb1fefc61833068880cf3546f217e4eb3

Observation 178320c8-2265-497d-8752-ec691fc40ac4 · inbound

Woosh: A Sound Effects Foundation Model cites this paper.

Woosh: A Sound Effects Foundation Model Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:53:15.813464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T20:51:08.144573Z digest=sha256:6130457c26ee6161a9c19faa86e1bc217a1f752252a9cfab0178d75afc5aa118

Observation a0c32f32-2706-4912-86a0-4f01aa865198 · inbound

UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations cites this paper.

UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:19:19.976408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T10:18:16.972414Z digest=sha256:98a4ba4c9f60ab5b54ccf7da589261588205a91d5135987fc07cb40beaf0b2f3

Observation 9a992810-63b1-45e6-81b4-95ddc09f16ff · inbound

UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations cites this paper.

UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T16:16:37.406485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:16:37.406485Z digest=sha256:ea5103e973a3dbe9d7decbc6c6f07f088e55f57d19eb76e8236774eec68d02d8

Observation 311e960d-e47a-4882-9ac5-9415515192ac · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.843462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:348b563af993f83273ae394880317c753c3bc351efc93b8e24436bd409e6ebc1

Observation 648201a4-36d9-4773-a494-e62e2afccb2f · inbound

BareWave: Waveform-Native Flow-Matching Text-to-Speech cites this paper.

BareWave: Waveform-Native Flow-Matching Text-to-Speech Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:35.201518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T15:14:03.681833Z digest=sha256:05f1cc57073415c7cad363b0b8f4e1cf9c40c9f0f44b5ca00ed265086a35e20f

Observation 99123b19-0579-4663-b5b5-e39575b762bb · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.635697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:d6b4f94aeee0b3156457e333304b57c2728f9462e4f7b075e734cca254df7212

Observation 6c09b5b2-5213-461a-bda2-ccf89aa39fb3 · inbound

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations cites this paper.

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:58:08.358364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T08:38:51.371500Z digest=sha256:781fda16b99d3c4e445d7d950d4f3b6a9ddfaa5e131a25f82a0b73c5f5f7efee

Observation 8dec23c5-03d7-4633-a10e-3d92de0d2bad · inbound

PhASE-Flow: Phonetic-Conditioned Acoustic Flow Matching in SSL Representation Domain for Speech Enhancement cites this paper.

PhASE-Flow: Phonetic-Conditioned Acoustic Flow Matching in SSL Representation Domain for Speech Enhancement Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:09:00.912005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T23:05:08.894347Z digest=sha256:d03b5b93e8a792c036bfcc763713ca3aee77df0fb987bb0ded4625bef0638841

Observation 432d4ae4-8932-40ba-b173-96e958d3901b · inbound

NAC: Neural Action Codec for Vision-Language-Action Models cites this paper.

NAC: Neural Action Codec for Vision-Language-Action Models Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:59:37.890291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T13:59:53.484305Z digest=sha256:f29cac8c881b534ad2e85e2795f4d05c0806a3527025d638989a0f4081648719

Observation 662649d3-2a3e-464e-aed4-30e1c9418328 · inbound

Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS cites this paper.

Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:29:51.743004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T06:49:45.405554Z digest=sha256:fee30641c065ba61f58d4fd20252bc6f05c5337e6f8f1bc7f0372dc28dcea2a1

Observation cd3727ec-abf2-4b56-94b8-836d6b00c7f6 · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:19:49.654721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T07:02:36.499424Z digest=sha256:9eb61db98b353d061e9a808858335772299c35528a1545f950a8252778f4476d

Observation 4c3a6dcd-7094-4446-ad5e-af2b5c4292ab · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T12:44:20.831164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:44:20.831164Z digest=sha256:175293d1b9b37f7c8a6334c71ea4bf9c64b429c08b1a755b8764cccb831af2d1

Observation 3d2018f4-9fdd-48aa-9fc2-f2f9f6ddbafc · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 116

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:55:42.189973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:88ec376b9d3f00353d5235416a79a1b6cf7c4e4fa28ff1a3012f8ce659324c85

Observation 898e4653-2033-4961-a5d5-9e5126652721 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 167

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.674104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:4a1c713ef58f542c9f6a83a63820436536292b652a76be90a1b632fb3cc0087f

Observation 824fa72a-15bd-48fd-ae54-090373c1d9e3 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 167

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:eba3e67f2e53e49d91db950dacbf81af2c81ece081ebdc0a636f7413a13272d4

Observation cf042f65-49d9-4778-b6aa-65d59d4f3b99 · inbound

ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching cites this paper.

ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T06:31:32.258567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:31:32.258567Z digest=sha256:c5c281d436451d72a3e395905188c82918e49bd9ba534d62d96505b849f95403

Observation 4922feff-a1cd-4b7e-97b4-377f7cebfe07 · inbound

Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer cites this paper.

Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T18:46:14.038134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:46:14.038134Z digest=sha256:79f661da77772f4646b0bf3ac5d68bda0338485d9abc6ccec439f2ab5ce85f69