{"work":{"id":"dfb28242-44a0-4f86-bb57-65dd1aae8947","openalex_id":"https://openalex.org/W4379259581","doi":"10.48550/arxiv.2306.00814","arxiv_id":"2306.00814","raw_key":null,"title":"Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis","authors":null,"authors_text":"H","year":2023,"venue":"cs.SD","abstract":"Recent advancements in neural vocoding are predominantly driven by Generative Adversarial Networks (GANs) operating in the time-domain. While effective, this approach neglects the inductive bias offered by time-frequency representations, resulting in reduntant and computionally-intensive upsampling operations. Fourier-based time-frequency representation is an appealing alternative, aligning more accurately with human auditory perception, and benefitting from well-established fast algorithms for its computation. Nevertheless, direct reconstruction of complex-valued spectrograms has been historically problematic, primarily due to phase recovery issues. This study seeks to close this gap by presenting Vocos, a new model that directly generates Fourier spectral coefficients. Vocos not only matches the state-of-the-art in audio quality, as demonstrated in our evaluations, but it also substantially improves computational efficiency, achieving an order of magnitude increase in speed compared to prevailing time-domain neural vocoding approaches. The source code and model weights have been open-sourced at https://github.com/gemelo-ai/vocos.","external_url":"https://arxiv.org/abs/2306.00814","cited_by_count":11,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2306.00814","created_at":"2026-05-10T10:19:19.908625+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Vocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis","render_title":"Vocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis"},"hub":{"state":{"work_id":"dfb28242-44a0-4f86-bb57-65dd1aae8947","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":15,"external_cited_by_count":11,"distinct_field_count":4,"first_pith_cited_at":"2024-10-09T13:46:34+00:00","last_pith_cited_at":"2026-07-06T15:11:57+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-23T20:39:49.668565+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":1},{"context_role":"method","n":1}],"polarity_counts":[{"context_polarity":"background","n":1},{"context_polarity":"use_method","n":1}],"runs":{},"summary":{},"graph":{},"authors":[]}}