Pith. sign in

Paper Citation Record · LEDGER

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments

As of 8 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2506.03554.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03554 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:03:09.237635Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:03:09.067600Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:03:09.354907Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact2
  • verified fuzzy38
  • unresolved5
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2efe9a14-4863-4878-84fe-2e09512cdfbe · outbound

This paper cites One key component of this progress is neural vocoders, which synthesize audio waveforms from acoustic fea- tures.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments One key component of this progress is neural vocoders, which synthesize audio waveforms from acoustic fea- tures

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.753919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.063905Z digest=sha256:548e2d286bb046fcd09eff7422a1a89ee72ade97f7bbcf2a9d6497564723509f

Observation a8644e48-5651-4605-8c30-96adb3581fc6 · outbound

This paper cites Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:03:09.358217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.067600Z digest=sha256:a3b599459cadfe1bda9d40c72f1e9322f868ca2915abc73c45406433d54d1d0f

Observation 8c077676-0d9f-4557-8c55-2aa96aa17995 · outbound

This paper cites throughput We analyze the relationship between latency and throughput via block streaming synthesis using several neural vocoders.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments throughput We analyze the relationship between latency and throughput via block streaming synthesis using several neural vocoders

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.747067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.070559Z digest=sha256:d9c8fe1df1046259f5919ff918892db1bae2589fd78cb595b4172eb5f1b523c3

Observation b598b463-06d1-4f0e-8c14-9b8e05cc4152 · outbound

This paper cites Wavehax and MS-Wavehax utilized F0 for generating prior signals, whereas other models concate- nate it with the mel-spectrogram, resulting in a 101-dimensional input feature.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Wavehax and MS-Wavehax utilized F0 for generating prior signals, whereas other models concate- nate it with the mel-spectrogram, resulting in a 101-dimensional input feature

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:03:09.739379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.073128Z digest=sha256:db7622602c3c403996c5f2fe80e8bfdc107bcc32a562bb026c5223492f048fdf

Observation 8a563839-47b6-4221-ad58-b0e62c571d16 · outbound

This paper cites Next, we evaluate its speech quality under causal and non-causal condi- tions, compared to the vocoders described in Section 3.1.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Next, we evaluate its speech quality under causal and non-causal condi- tions, compared to the vocoders described in Section 3.1

Reference 5

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:03:09.731975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.075967Z digest=sha256:e823e48b1ba8538d022078ba10f172e0e461a97fc513b3729ea9ac161b2019ee

Observation 494c870b-c8e6-48d0-ac85-86a610e4cf49 · outbound

This paper cites an unresolved cited work.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:03:09.704463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.084068Z digest=sha256:222f084cf9a90765c7c2f0bf4e9b0ca94494e3b983d6edba62f46f159479ceed

Observation 13749803-0ece-4dba-907e-46532bc52575 · outbound

This paper cites Our analysis revealed that streaming throughput de- pends on overhead from data and parameter loading as well as computational complexity.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Our analysis revealed that streaming throughput de- pends on overhead from data and parameter loading as well as computational complexity

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.714250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.081498Z digest=sha256:614da706f1c6c4a4ec522a1f1cfbe588d454d22186458ba9471e4c03865d8f9e

Observation 740076dc-5257-4382-91a1-d16bf60bde88 · outbound

This paper cites Generative Ad- versarial Nets,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Generative Ad- versarial Nets,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.694699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.086457Z digest=sha256:df64dee72de345c1cd0c4a4c177bc4f9e8f089b815b6aafe331257abe2b436c4

Observation 1fa39c74-ec56-4793-a3e2-f68e742d76e0 · outbound

This paper cites MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments MelGAN: Generative Adversarial Networks for Conditional Waveform Synthesis,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.685081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.089020Z digest=sha256:e1f721217ea656ecb6c8a05c807f0e8a7c2baf1ab0fa8375cba9841c1a78bd11

Observation a7037021-c55d-4037-ae5e-bf7ddcf94ce2 · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.675221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.091510Z digest=sha256:b8b83ae8ad6d3492fd7c9217a8764d3cdb6240352218a9ae332e4b4e49685f78

Observation f31581db-3072-41f3-9d29-a576c8f35400 · outbound

This paper cites iSTFTNet: Fast and Lightweight Mel-Spectrogram V ocoder Incorporating Inverse Short-Time Fourier Transform,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments iSTFTNet: Fast and Lightweight Mel-Spectrogram V ocoder Incorporating Inverse Short-Time Fourier Transform,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.664184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.095332Z digest=sha256:25f379971e94cb26b031cbfde974859cf178f76605bd4b1c5bbf6a955f77b2a8

Observation d080a97c-771d-4f78-a9d7-2ee0952f5583 · outbound

This paper cites iSTFTNet2: Faster and More Lightweight iSTFT-Based Neural V ocoder Using 1D- 2D CNN,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments iSTFTNet2: Faster and More Lightweight iSTFT-Based Neural V ocoder Using 1D- 2D CNN,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.654149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.098995Z digest=sha256:cc442949f9b1ed7d059d00dfeed0ca3fd2a41448dea8994aa327b9b6c88d2602

Observation 8574965a-ee7c-4caa-b7cf-233f52f36f98 · outbound

This paper cites APNet: An All-Frame-Level Neural V ocoder Incorporating Direct Prediction of Amplitude and Phase Spectra,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments APNet: An All-Frame-Level Neural V ocoder Incorporating Direct Prediction of Amplitude and Phase Spectra,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.644340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.102904Z digest=sha256:2ec52bc483430eef47117d01032453199b2be04ed45bfbfd73c8eaafe538a27c

Observation ba0011ed-c82a-492f-b77c-94be19529883 · outbound

This paper cites V ocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments V ocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.634797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.106552Z digest=sha256:55ad574a42e156832c56784ef1b1aae75dc8136a6fd23d0d8a87eef728127b3c

Observation 05937b8d-3552-4d4a-97ab-122d250385ef · outbound

This paper cites AC-VC: Non-Parallel Low La- tency Phonetic Posteriorgrams Based V oice Conversion,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments AC-VC: Non-Parallel Low La- tency Phonetic Posteriorgrams Based V oice Conversion,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.574398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.140221Z digest=sha256:633207ddbb44a6aa281fe7a027f6e2b2bb5cb11234c1aae1dff6aff20a600104

Observation dcec3a08-34b7-45bf-a098-15219e497a6e · outbound

This paper cites Low-latency real-time non-parallel voice conversion based on cyclic variational autoencoder and multiband WaveRNN with data-driven linear prediction,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Low-latency real-time non-parallel voice conversion based on cyclic variational autoencoder and multiband WaveRNN with data-driven linear prediction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.625084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.113880Z digest=sha256:440caf1821252b0041c97894e7a8be8feb81b1d4f92d816661e3b0f1848c2183

Observation cad249c8-ffe8-45be-a510-2acefc71d38f · outbound

This paper cites An Investigation of Streaming Non-Autoregressive sequence-to-sequence V oice Con- version,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments An Investigation of Streaming Non-Autoregressive sequence-to-sequence V oice Con- version,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.615267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.117576Z digest=sha256:ab2827f04d84abc2411ebe0fca1193b9264962fef0d197ebc7d1bc68241abce5

Observation 6ed446e1-84ab-4cac-9979-056845733a0b · outbound

This paper cites Streaming non-autoregressive model for any-to-many voice conversion.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Streaming non-autoregressive model for any-to-many voice conversion

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:03:09.347294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.121069Z digest=sha256:750fc0d0ac8aa2fdf441dd8a06258f3634656022611f4f16996b4e4c53465e54

Observation a06f59ae-1390-4505-ac30-b3f3a3e5baa3 · outbound

This paper cites Wavehax: Aliasing-Free Neural Waveform Synthesis Based on 2D Convo- lution and Harmonic Prior for Reliable Complex Spectrogram Es- timation,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Wavehax: Aliasing-Free Neural Waveform Synthesis Based on 2D Convo- lution and Harmonic Prior for Reliable Complex Spectrogram Es- timation,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:09.124691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:09.124691Z digest=sha256:464100a0f876a5ee59d1708dfa17f3672227f361046b6bf8acefe1300448f824

Observation 374ba3cc-00c5-42de-9e55-74dce93c86fe · outbound

This paper cites Multi-Stream HiFi-GAN with Data-Driven Waveform Decomposition,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Multi-Stream HiFi-GAN with Data-Driven Waveform Decomposition,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.605762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.128505Z digest=sha256:231b64c810871a38014548e2f4a3d1787421266202bd766e4b770a2f066385d5

Observation 6dd36464-7ac6-40f8-b187-9d9a403187c4 · outbound

This paper cites Implementation of DNN-based real-time voice conversion and its improvements by audio data augmentation and mask-shaped device,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Implementation of DNN-based real-time voice conversion and its improvements by audio data augmentation and mask-shaped device,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.595860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.132327Z digest=sha256:36ba759be47262146b16939e639635fd1cd60c8618b5b43d92475df731690cde

Observation edb05948-0859-4b18-ab05-6af0f6f97abd · outbound

This paper cites Real-Time, Full-Band, Online DNN-Based V oice Conversion System Using a Single CPU,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Real-Time, Full-Band, Online DNN-Based V oice Conversion System Using a Single CPU,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.585450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.136347Z digest=sha256:c2754c77010c17fa43deefbb321561db250086a3ccfd26f71a42713fcb54f4be

Observation 21d67325-1765-4f3a-ace1-81d526740d34 · outbound

This paper cites Fregrad: Lightweight and Fast Frequency-Aware Diffusion V ocoder,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Fregrad: Lightweight and Fast Frequency-Aware Diffusion V ocoder,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.495782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.169405Z digest=sha256:52e408b874a7894f732126f01e2ee1fc416ab940e27c0cac925964c59086e980

Observation 1e4c5a7c-83fe-45ba-8011-8ee975ddc3de · outbound

This paper cites Incremental Text-to-Speech Syn- thesis with Prefix-to-Prefix Framework,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Incremental Text-to-Speech Syn- thesis with Prefix-to-Prefix Framework,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.563771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.144149Z digest=sha256:26973b30021ce4bfdd7213275ff48d7ebcf266a4e8bf0e10c2101bbc93db47c7

Observation 6adba2ad-acc7-4c00-8f9e-fcfce809fe87 · outbound

This paper cites Neural iTTS: Toward Synthesizing Speech in Real-time with End-to-end Neural Text- to-Speech Framework,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Neural iTTS: Toward Synthesizing Speech in Real-time with End-to-end Neural Text- to-Speech Framework,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.554309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.147888Z digest=sha256:12789666eeed43d6d3470f2ef7c2453ad827681b1464f6c925a616dcf87c271b

Observation 8cf19f83-7642-40c9-a31d-c81d3184e6f2 · outbound

This paper cites High Qual- ity Streaming Speech Synthesis with Low, Sentence-Length- Independent Latency,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments High Qual- ity Streaming Speech Synthesis with Low, Sentence-Length- Independent Latency,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.544086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.151384Z digest=sha256:8af69970d493bd47b24cc8ba1841bec85eb3d9a05c1de66acb49c0e97dafd02a

Observation f573e596-7e4c-4d92-942e-0198c6822814 · outbound

This paper cites A ConvNet for the 2020s,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments A ConvNet for the 2020s,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.534826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.155362Z digest=sha256:205d44d316b19c103a0ac138479889f26fa36653ec94c967f0e7022531c1abfc

Observation 58d7f9f8-d176-4597-b13b-69bfbb865c2b · outbound

This paper cites Design and evaluation of parallel quadrature mirror filters (PQMF),.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Design and evaluation of parallel quadrature mirror filters (PQMF),

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.524902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.159040Z digest=sha256:831678dc07b042c6f7af72c03b496f214e0575b7529830bf57605b3aaa8a6c51

Observation 79ddeef0-150f-4f4b-8d6b-192b688b8db6 · outbound

This paper cites Multi-band MelGAN: Faster Waveform Generation for High-Quality Text-to-Speech,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Multi-band MelGAN: Faster Waveform Generation for High-Quality Text-to-Speech,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.514584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.162515Z digest=sha256:cbb59ceaf8a04f1ecb85d426ca12646266bfcb7e1139dd9ad5be0783b67f003a

Observation 4a2daaeb-b25f-4cfd-ae22-8b1c61441a2b · outbound

This paper cites Fre-GAN: Adversar- ial Frequency-Consistent Audio Synthesis,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Fre-GAN: Adversar- ial Frequency-Consistent Audio Synthesis,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.505221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.165700Z digest=sha256:49c1fa4f1ca3d631974e8a80c93955b209a9d187ae69fe6292e10f3c2c924a71

Observation 6edea2f9-de32-41e3-b0d2-d48901f5b0bf · outbound

This paper cites Harvest: A High-Performance Fundamental Fre- quency Estimator from Speech Signals,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Harvest: A High-Performance Fundamental Fre- quency Estimator from Speech Signals,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.438612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.202363Z digest=sha256:07d5bd40446c5f9a8155b28ba2e5126ab524ab7e0db4aad7b31f5e290f337565

Observation dbbbdaa0-e510-4d26-9275-092f49970bb8 · outbound

This paper cites Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assess- ment of telephone networks and codecs,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assess- ment of telephone networks and codecs,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.486771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.175261Z digest=sha256:da82d67a621c95b7a2b6c860bba23480a356bad70b7e558288100e611a9a6eed

Observation d1992397-2144-459f-bdb1-fbcb3c816a92 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for V oiceMOS Challenge 2022,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments UTMOS: UTokyo-SaruLab System for V oiceMOS Challenge 2022,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.476818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.179126Z digest=sha256:0eccad2c1d239feb98f70c19e4beda6c2779c20f9b23e1a10d3649760e67ee00

Observation 4269562f-f098-431a-b358-b1c06ec36040 · outbound

This paper cites Lightweight and High-Fidelity End-to-End Text-to-Speech with Multi-Band Generation and Inverse Short-Time Fourier Transform,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Lightweight and High-Fidelity End-to-End Text-to-Speech with Multi-Band Generation and Inverse Short-Time Fourier Transform,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.467769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.183017Z digest=sha256:04631fca8aaa8120f513214731457bb7e91b674663790a9f93a72b420e00e9f4

Observation bfbb5934-af47-45e0-97ac-8b70ef0a9ed6 · outbound

This paper cites Aggregate.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Aggregate

Reference 36

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:03:09.723585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.078818Z digest=sha256:e11e3248e2b4c8fd3d309d060774cda9fc720a791c210183b5a1c6d9e570a054

Observation 627f44c1-21ec-4db1-b577-10cc0ac60fb8 · outbound

This paper cites Developing Real-Time Stream- ing Transformer Transducer for Speech Recognition on Large- Scale Dataset,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Developing Real-Time Stream- ing Transformer Transducer for Speech Recognition on Large- Scale Dataset,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.458455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.186530Z digest=sha256:e9eebd93a7d928c094eb67c8065b874a6d2b3e07c6df649007aa464cbc4a684e

Observation de3f64f1-3a9d-4411-a7e0-03fffb896ebe · outbound

This paper cites Layer Normalization.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Layer Normalization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:09.190271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:09.190271Z digest=sha256:9bff6a2c9d2dd62b74634d33a5ace37cc76a405a328cf017a180e3e479f5e9e8

Observation 98d3165d-2e01-430b-a263-ff9dd63d3cf4 · outbound

This paper cites Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.448743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.194351Z digest=sha256:6b87fad0d1813761b547dec6d91ce6e8b56fdb47757791b497d4362278a4cc21

Observation cd0c716b-1b79-40bb-9899-f87c7b297c54 · outbound

This paper cites JVS corpus: free Japanese multi-speaker voice corpus.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments JVS corpus: free Japanese multi-speaker voice corpus

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:09.198316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:09.198316Z digest=sha256:c14f80c110c2cbf0d3bc6ba4bb8127e61a471ebf6427ee830fb7a6ffa7c1fb5d

Observation 4b513dda-1176-4d8c-ab2c-e07f04eb3d9d · outbound

This paper cites BigVGAN: A Universal Neural V ocoder with Large-Scale Training,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments BigVGAN: A Universal Neural V ocoder with Large-Scale Training,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.428879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.206313Z digest=sha256:7def01f63e3942b19cced718c31f5f52152104353eebb23df8eb759d290ff932

Observation 5128eaf3-493a-497b-9c95-2a6f70845f3a · outbound

This paper cites UnivNet: A Neural V ocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments UnivNet: A Neural V ocoder with Multi-Resolution Spectrogram Discriminators for High-Fidelity Waveform Generation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.418774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.210135Z digest=sha256:196bf4a254064be55fe5f95bcce1f2ce2c31851f6afdef091f4df0ce4341a48d

Observation 467c237e-279d-48b1-87dd-b5a5fb64ba68 · outbound

This paper cites JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments JSUT corpus: free large-scale Japanese speech corpus for end-to-end speech synthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:03:09.213625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:03:09.213625Z digest=sha256:d452d62250aed064e10a357e9018e721e90bd79dd1ea2750695adcadf2c44811

Observation 9c145dd8-8aec-461c-b1cf-893d1d63cae1 · outbound

This paper cites Matcha-TTS: A Fast TTS Architecture with Conditional Flow Matching,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Matcha-TTS: A Fast TTS Architecture with Conditional Flow Matching,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.408771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.217924Z digest=sha256:04aeb788f45a2898042dc8d7a77086ed419f143d2e8f66a429e1ec91b1d6ead9

Observation 81d8bd2e-2a86-4167-b268-d0ff0e89b1a6 · outbound

This paper cites Attention is All you Need,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Attention is All you Need,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.400038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.221656Z digest=sha256:fc3fd843dce8e154f0fd5167f5fbe981086c9911fdba9a9700c6056a73024eea

Observation a7c8b92f-06ea-41dc-a53e-f2366fc7cbe3 · outbound

This paper cites Flow Matching for Generative Modeling,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Flow Matching for Generative Modeling,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.392214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.225535Z digest=sha256:c4322eeeedefda80d536116584a31a4228488796c9ea85e839e05b996318952a

Observation 567c61dd-d2d7-4d1b-99f0-b6810a1a0687 · outbound

This paper cites Matcha-tts.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Matcha-tts

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.382954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.229084Z digest=sha256:f2da8e9fc829fd8d71fc9e507f70a8bd464c7a36cbccd609bdd2512ae36e6ac3

Observation 7d0d27f3-920b-440a-b8bc-65fd66d3f718 · outbound

This paper cites jsut-label.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments jsut-label

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.374466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.232821Z digest=sha256:0b1d12463bd4601dd8a4870814cfc979c204763f85a7d267e116d7f346ccb8a6

Observation 17222672-317a-478c-8678-94b2f5604230 · outbound

This paper cites What the Future Brings: Investigating the Impact of Lookahead for Incremental Neural TTS,.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments What the Future Brings: Investigating the Impact of Lookahead for Incremental Neural TTS,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:03:09.366129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.237635Z digest=sha256:14695aad4f2ad2c55135135cf8ed00d4adc9023c6eea00735b8cc51e29ddce60

Pith citing papers

Observation a8644e48-5651-4605-8c30-96adb3581fc6 · inbound

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments cites this paper.

Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:03:09.358217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:03:09.067600Z digest=sha256:a3b599459cadfe1bda9d40c72f1e9322f868ca2915abc73c45406433d54d1d0f