Pith. sign in

Paper Citation Record · LEDGER

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram

As of 15 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 0 inbound Pith citation observations for arXiv:2411.11258.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11258 v1

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T18:47:51.689848Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

73 of 73 outbound references displayed

  • verified exact1
  • verified fuzzy53
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c47f90a2-7987-4b51-b815-cbf24766bb4a · outbound

This paper cites APCodec: A Neural Audio Codec with Parallel Amplitude and Phase Spectrum Encoding and Decoding.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram APCodec: A Neural Audio Codec with Parallel Amplitude and Phase Spectrum Encoding and Decoding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.288451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.288451Z digest=sha256:f5fb8d8b10bfc80e9e894af22b455ff88f532cbbd70283783153785cea5ac7e0

Observation 4b56eaca-498e-458e-ab90-63622f40485e · outbound

This paper cites an unresolved cited work.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:47:53.122993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.295087Z digest=sha256:95972cb812034945e785b7f8380c5e3247995da0a0769859c3da6f029f545563

Observation e5a4ab15-641f-4883-a496-1466aaef0844 · outbound

This paper cites APNet: An all-frame-level neural vocoder incorpo- rating direct prediction of amplitude and phase spectra.IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram APNet: An all-frame-level neural vocoder incorpo- rating direct prediction of amplitude and phase spectra.IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:53.104762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.300069Z digest=sha256:a7cbf9750d8bb138b0dd8fb398f35d8806f8c6afce5541995c6977008f9ad495

Observation 941397a5-4586-4567-b8d9-13954a1f2a18 · outbound

This paper cites Neural speech phase prediction based on parallel estimation architecture and anti-wrapping losses.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Neural speech phase prediction based on parallel estimation architecture and anti-wrapping losses

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:53.089020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.305152Z digest=sha256:5a03a9943cf987bd344a3f07790c7a20104aab12228dbf56caa9c43f5ef61b3c

Observation 8acb7ca0-f313-455d-9cff-e87352053c20 · outbound

This paper cites Long-frame-shift neural speech phase prediction with spectral continuity enhancement and interpolation error compen- sation.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Long-frame-shift neural speech phase prediction with spectral continuity enhancement and interpolation error compen- sation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:53.071416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.310478Z digest=sha256:e537d442e95ca24d3ec614fd50ecce89bf11e86bac8fac8b76c8e892f036cd66

Observation 1ed81267-acbe-47a6-b11a-49bbf046eafd · outbound

This paper cites vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.315555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.315555Z digest=sha256:f6a79cbad991f35733da04079ac147d3fb7fe3336d92a6d2431166dce53cd1ae

Observation 697d32cd-6f82-431f-8350-602e496b66b5 · outbound

This paper cites Audiolm: a language modeling approach to audio generation.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Audiolm: a language modeling approach to audio generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:53.053852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.321616Z digest=sha256:9279d92f7b6bf04318d220949ca57e745cc2169768bc8b2c57c061abb6f72ce8

Observation 9bbf6214-0846-452c-8ca2-5aa583c8fc1b · outbound

This paper cites Iso/mpeg-1 audio: A generic standard for coding of high-quality digital audio.Journal of the Audio Engineering Society , 42(10):780–792, 1994.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Iso/mpeg-1 audio: A generic standard for coding of high-quality digital audio.Journal of the Audio Engineering Society , 42(10):780–792, 1994

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:53.032029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.326911Z digest=sha256:b024aa5c794575af96983185a80812adba954b7d3e20f2a83c0e21fb8d4b6b13

Observation f992542b-b0d6-4b1a-801a-b6db205cd865 · outbound

This paper cites Crowdsourcing preference tests, and how to detect cheating.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Crowdsourcing preference tests, and how to detect cheating

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:53.014682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.331806Z digest=sha256:9a5d25e28514fd5f7b959c15e08257ed16697a41f92283c9540e67b79741c9de

Observation 3cb54547-c291-4b60-aed2-294195bd0fe2 · outbound

This paper cites ViSQOL v3: An open source production ready objective speech and audio metric.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram ViSQOL v3: An open source production ready objective speech and audio metric

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.997282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.337005Z digest=sha256:fc589d27344f79690ba408c83a804a3f0823e654264bdb80207c11d394eab61a

Observation 96a17bfa-ada3-48f8-ace9-120a324aa15e · outbound

This paper cites High fidelity neural audio compression.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram High fidelity neural audio compression

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.980738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.342222Z digest=sha256:7b35157e53f08e581ba0b2a89864f19c555b63550723e500674fd4580a65aff1

Observation 2b2e2821-aa0f-491d-8f20-500d520096b7 · outbound

This paper cites Overview of the evs codec architecture.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Overview of the evs codec architecture

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.964100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.347324Z digest=sha256:4be7f7a07c55402891fed1a0fd4951770b1430694c2209dbb8fd7477d8fc20da

Observation 76fcf08f-aaab-403a-9db2-d61b6e5653d0 · outbound

This paper cites Adversarial audio synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Adversarial audio synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.352694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.352694Z digest=sha256:174a1771153e1792719a990afeda271d31bff4de584f1064b12a843cb1fba6b3

Observation 3a511ba9-fbc4-4dca-870e-1e75faf706da · outbound

This paper cites VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram VQTTS: High-Fidelity Text-to-Speech Synthesis with Self-Supervised VQ Acoustic Feature

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.357279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.357279Z digest=sha256:4bb94dda42668986740d0227f3bab0f50aa81b310113147484e90abd332ccffa

Observation ed038733-e9fb-42d2-b396-b4060d4c58c0 · outbound

This paper cites Apnet2: High-quality and high-efficiency neural vocoder with direct prediction of amplitude and phase spectra.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Apnet2: High-quality and high-efficiency neural vocoder with direct prediction of amplitude and phase spectra

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.937121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.363643Z digest=sha256:7fbffd148215eedd51f6c381bc4258b864e6f506c03c6534d3187d4bfd34ad16

Observation d6403e04-e83a-4a2c-992e-ab1a71551444 · outbound

This paper cites Generative adversarial nets.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Generative adversarial nets

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.921959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.369873Z digest=sha256:7e054b695c8a1a7b38089141f9c6f291e995572a89c3eb8f978910ab35850bf2

Observation c04a2b58-0840-49c9-b290-0f5e592b06a0 · outbound

This paper cites Long short-term memory.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Long short-term memory

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.905221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.375651Z digest=sha256:055496371a734ea2911e5da5354d2a32f60d713e8479c6dc32e0f383de06e74b

Observation 4865663c-5a24-4fed-8618-85e055c53f84 · outbound

This paper cites A spectral energy distance for parallel speech synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram A spectral energy distance for parallel speech synthesis

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.889445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.381364Z digest=sha256:cece1d64712a2da393b92c65b8200112fa24e2b35c410c098389cb55f9ded44a

Observation c6074c2d-65be-49b0-aeb0-ebc65e1dcb41 · outbound

This paper cites A Multi-Stage Multi-Codebook VQ-VAE Approach to High-Performance Neural TTS.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram A Multi-Stage Multi-Codebook VQ-VAE Approach to High-Performance Neural TTS

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-12T18:47:51.902113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.386579Z digest=sha256:648a455df418335298a9f045f3f4adb54c70e93ae75c0fa74ef2af1c8d37d5e4

Observation d03e39e0-be2c-4a5f-9454-7870c72894c8 · outbound

This paper cites Deep residual learning for image recognition.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Deep residual learning for image recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.393319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.393319Z digest=sha256:43e79797ded9c8827182976ca0e62e048f8fb3c9be4ca0ab1ccb299be41d3b3c

Observation 55f6247c-120e-4e43-8e06-93bea41fbc63 · outbound

This paper cites Gaussian error linear units (GELUs).

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Gaussian error linear units (GELUs)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.862310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.398742Z digest=sha256:e4eaba4101e9ef25e3dded14aa30a804c12e30244b28f670fd5ad6402abe2e55

Observation a225fc0b-6077-4c5f-9413-b594abc6ca26 · outbound

This paper cites Hubert: Self-supervised speech repre- sentation learning by masked prediction of hidden units.IEEE/ACM Transactions on Audio, Speech, and Language Processing , 29:3451–3460, 2021.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Hubert: Self-supervised speech repre- sentation learning by masked prediction of hidden units.IEEE/ACM Transactions on Audio, Speech, and Language Processing , 29:3451–3460, 2021

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.845446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.403526Z digest=sha256:20700c6a5aa263c3c79a509091754be5e229836a6490ec5f0d11cbe960cfa689

Observation 220e17b8-b852-496c-9776-fb4f4a0d4001 · outbound

This paper cites RepCodec: A Speech Representation Codec for Speech Tokenization.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram RepCodec: A Speech Representation Codec for Speech Tokenization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.409362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.409362Z digest=sha256:f0c0582144a63b5a01314a617c2e2cd4b0fa61903409f5e31af92afa6ad47e0d

Observation 194898d2-e1b8-4824-9575-de9a4ccab147 · outbound

This paper cites The LJ speech dataset.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram The LJ speech dataset

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.829906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.414859Z digest=sha256:5f24501c1a0a1597b9ca69516e03abb6199d2d444d5674a6094a12ac3788a487

Observation 5843b434-5ecb-4fc7-91d5-ce407d34d7eb · outbound

This paper cites Univnet:Aneuralvocoder with multi-resolution spectrogram discriminators for high-fidelity waveform gener- ation.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Univnet:Aneuralvocoder with multi-resolution spectrogram discriminators for high-fidelity waveform gener- ation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.814549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.419681Z digest=sha256:529a04f8a61f474c466533738a5f968baf870dd81fbb8a1fb93541ffd8051649

Observation 0f21d5a1-6e49-49fc-a94a-538caa4c5177 · outbound

This paper cites GlotNet—a raw waveform model for the glottal excitation in statistical parametric speech synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram GlotNet—a raw waveform model for the glottal excitation in statistical parametric speech synthesis

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.798943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.425218Z digest=sha256:ec291affc67268e02b6a63cf52ca57cc3253c447e100eacc68f354126f7125d2

Observation b4063528-f205-44e2-b288-744ca791cc66 · outbound

This paper cites iSTFTNet: Fast and lightweight mel-spectrogram vocoder incorporating inverse short-time Fourier transform.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram iSTFTNet: Fast and lightweight mel-spectrogram vocoder incorporating inverse short-time Fourier transform

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.783108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.430599Z digest=sha256:f77bcc5a674fb960a7565fbb4eba21437bf9d3551c9f45dcf399155129f375cf

Observation ccc19170-6be2-4ceb-8727-b9059f02950c · outbound

This paper cites an unresolved cited work.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:47:52.764666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.437083Z digest=sha256:54961f6ced715a4a5fed6c2b810d99ad8d69afc34a464ce3771358872fab4d3a

Observation 499348de-367a-4e3d-b0c0-1b7306d3e7cb · outbound

This paper cites HiFi-GAN: Generative adver- sarial networks for efficient and high fidelity speech synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram HiFi-GAN: Generative adver- sarial networks for efficient and high fidelity speech synthesis

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.748534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.442900Z digest=sha256:c31ac05c28d45480ce96e79150054beec40fbdee0b74040f60479b9978fad1bc

Observation a5ab773e-2581-48cb-a717-aa434258e7e0 · outbound

This paper cites an unresolved cited work.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:47:52.731393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.448372Z digest=sha256:ce427a84a300764aeabc544d5c4bb21f5ef5a5c429e3970948fcf73a60830a66

Observation 4fdfc843-8109-4659-ab48-0fae4904b63d · outbound

This paper cites MelGAN: generative adversarial networks for conditional waveform synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram MelGAN: generative adversarial networks for conditional waveform synthesis

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.715308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.453836Z digest=sha256:79e876ffa32c260e8b8703dc50b85d38340499e5965c7e86dc8e04118fc0ee3f

Observation cce4f8fb-943f-431b-9f79-c3c1df243fa0 · outbound

This paper cites PHASEAUG: A differentiable augmentation for speech synthesis to simulate one-to-many mapping.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram PHASEAUG: A differentiable augmentation for speech synthesis to simulate one-to-many mapping

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.699637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.459021Z digest=sha256:d480b7721bb29feec5281198ce9eba13327bae7b381e8358b3bd449ba73be0ae

Observation 430f0ddd-fbfb-465c-921b-6aaaf1760bc2 · outbound

This paper cites Neural vocoder is all you need for speech super-resolution, 2022.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Neural vocoder is all you need for speech super-resolution, 2022

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.683429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.464048Z digest=sha256:52967ddc6a05aa6e4e54a4ba994dfd91dee09ca471daebc34d976310048f716c

Observation 7cc61c9e-f6aa-4b7b-bfcf-c4d1f4a034e3 · outbound

This paper cites DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram DiffGAN-TTS: High-Fidelity and Efficient Text-to-Speech with Denoising Diffusion GANs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.471319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.471319Z digest=sha256:42b6a2809846009dffcec021c82985ebd739f8fbbb9f88e2140b6ce7b9df50e8

Observation bf4bcccb-b280-4870-ac60-957f1ff73308 · outbound

This paper cites A convnet for the 2020s.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram A convnet for the 2020s

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.666050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.477762Z digest=sha256:d78f168ea8176d0f847364a15bb49225dfa7bcb46ba0db1c734a6c49b0572515

Observation 2653fa72-aaf2-41f2-9b47-7d3f33b439a8 · outbound

This paper cites Decoupled weight decay regularization.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Decoupled weight decay regularization

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.649809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.483846Z digest=sha256:0b209be1792c295d986830ff205086623d31fb8ce3b71d60eb6cec374f3d5b66

Observation ec25e984-31ea-416e-a5ea-549839a9ae5b · outbound

This paper cites Source-filter-based generative adversarial neural vocoder for high fidelity speech synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Source-filter-based generative adversarial neural vocoder for high fidelity speech synthesis

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.632295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.489915Z digest=sha256:b574896008b515e2b18d5da4df5dab4eead83f7290ace4d918f6e5db04a55b2f

Observation 8b4c81a1-b1bc-4688-99b2-95a34b50ec34 · outbound

This paper cites MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.494938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.494938Z digest=sha256:29188ddbb98a0a32077c923931f92cc90010f373993d5d9dcd5432396518fe31

Observation c2e6fa0c-9703-4bbc-93e8-b54035ab2da8 · outbound

This paper cites Rectifier nonlinearities improve neural network acoustic models.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Rectifier nonlinearities improve neural network acoustic models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.614586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.500806Z digest=sha256:983ad69db3f66d0d3e687f2b1bf70a066f9c01f21be4023daa7b9575fc65deb6

Observation 28320407-d618-4955-9d2c-260111eaf5c5 · outbound

This paper cites SampleRNN: An unconditional end-to-end neural audio generation model.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram SampleRNN: An unconditional end-to-end neural audio generation model

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.598385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.505578Z digest=sha256:d42db28e27710da30913d8086299c55fb4336370913902e7918368618cd66992

Observation a8c7ddc5-5a83-458b-81dd-b176c47734f2 · outbound

This paper cites Finite scalar quantization: Vq-vae made simple.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Finite scalar quantization: Vq-vae made simple

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.580994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.510632Z digest=sha256:c148c17435f079a0c242190bb1c30175524c5bf0c51f5cc1270dfa6bf573b01a

Observation c4024053-0c4c-4cc2-bc87-1602d3a5a2aa · outbound

This paper cites WORLD: A vocoder-based high-quality speech synthesis system for real-time applications.IEICE Transac- tions on Information and Systems , 99(7):1877–1884, 2016.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram WORLD: A vocoder-based high-quality speech synthesis system for real-time applications.IEICE Transac- tions on Information and Systems , 99(7):1877–1884, 2016

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.562869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.515933Z digest=sha256:1326cb4b21ba4700a318ca53402fdcc17b68753e485c1fab9b371549cc458e90

Observation c421bbc6-cd07-49eb-b690-2329909c9f50 · outbound

This paper cites Expediting tts synthesis with adversarial vocoding.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Expediting tts synthesis with adversarial vocoding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.542031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.521474Z digest=sha256:fcca688e8c2b4e730ff8b539688e55dd681c930d7a442bd4b0ab254ae50268c1

Observation a2a8e21e-e2b5-44d8-babd-9337ca96c34f · outbound

This paper cites EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.527155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.527155Z digest=sha256:6573e0b574c7b0e9a2049fbff6e688d488b1a7836af5269ff6bbc11055b70574

Observation 446f2159-383d-447f-9972-00152859cf80 · outbound

This paper cites Parallel WaveNet: Fast high-fidelity speech synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Parallel WaveNet: Fast high-fidelity speech synthesis

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.525246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.534346Z digest=sha256:d84c38fbd5df88f3bcfbc1493b17ba601ae96a0793b3b2945173434b6db1460b

Observation 846bffcc-0a1a-4ed1-8977-3dbd2776866c · outbound

This paper cites WaveNet: A generative model for raw audio.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram WaveNet: A generative model for raw audio

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.507049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.540109Z digest=sha256:946ce51e77a439260c82611c10ced7ecae83b348ddab36eb309e0882d9ea8351

Observation f1ca97b6-c3cf-4cdf-9b5b-4aaeba487477 · outbound

This paper cites Linear predictive coding.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Linear predictive coding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.489617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.544982Z digest=sha256:c563caab1c7df9f16a6f6bb6c4a239971a062d1d9d50b3e887c47cad0d078f20

Observation ee2e052b-9e6d-4b35-b7e3-f2e102498c0f · outbound

This paper cites Generativeadversarialnetwork-basedapproachtosignal reconstruction from magnitude spectrogram.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Generativeadversarialnetwork-basedapproachtosignal reconstruction from magnitude spectrogram

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.472077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.549935Z digest=sha256:3d91069314eb0075b5a002be9b70932f7a548eafd2a817687033242295363580

Observation 88d59813-fc92-40cf-8001-f5b6c91ff54a · outbound

This paper cites Wave-gan: a deep learning approach for the prediction of nonlinear regular wave loads and run-up on a fixed cylinder.Coastal Engineering, 167:103902, 2021.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Wave-gan: a deep learning approach for the prediction of nonlinear regular wave loads and run-up on a fixed cylinder.Coastal Engineering, 167:103902, 2021

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.452183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.554746Z digest=sha256:f2b4643b62d4e1139264ddb4fb2b13ee380d47b771519ecf405dad885d16e264

Observation 1cefcd94-0c6d-4474-bd30-1d4e6dffccf9 · outbound

This paper cites ClariNet: Parallel wave generation in end-to-end text-to-speech.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram ClariNet: Parallel wave generation in end-to-end text-to-speech

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.433624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.559532Z digest=sha256:e0f66d8c01a1d8113160d814d0cd7716f5a180011a855eb9789524a4b2aff7cc

Observation 8e4a7480-60ac-4177-892d-27b0e8a90bbc · outbound

This paper cites Waveflow: A compact flow- based model for raw audio.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Waveflow: A compact flow- based model for raw audio

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.406461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.564487Z digest=sha256:e10eac0824ffa0d35a5bdd56ca69d1c3976d715be317a31ca5c9b25b64dbd995

Observation a57b22a4-2418-4bee-8c32-e4bb6897a897 · outbound

This paper cites Waveglow: A flow-based gen- erative network for speech synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Waveglow: A flow-based gen- erative network for speech synthesis

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.389077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.570862Z digest=sha256:34468607c1bb8616d28a9e6e757f57436b6bae4af91d470787ac3c3f1331b41d

Observation 0ab47423-2800-45be-9f89-3af3eb931652 · outbound

This paper cites an unresolved cited work.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:47:52.372286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.576285Z digest=sha256:9661a68ceaf6f37c7be56bcc2458a8f6c585fd8ddc4e9859bba8085a406a7174

Observation 4a9d6b56-0a72-4a98-acf0-3178a7d499ff · outbound

This paper cites Fastspeech 2: Fast and high-quality end-to-end text to speech.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Fastspeech 2: Fast and high-quality end-to-end text to speech

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.356217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.582351Z digest=sha256:5da1736c6d357db660db675b5361cd940e9edd485846d7ff2d6d7a2bea389cde

Observation cce7f17d-e7df-428f-9ddd-745085f93a16 · outbound

This paper cites Fewer-token neural speech codec with time-invariant codes.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Fewer-token neural speech codec with time-invariant codes

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.338323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.587468Z digest=sha256:a6fe843941ecd51dd5ee1a18ab742d626ecdd0d2bbdcea2b39b18a2d6c10dfc7

Observation f35e040f-1fb3-4a21-bda6-ba0a552839fb · outbound

This paper cites Utmos: Utokyo-sarulab system for voicemos challenge 2022.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Utmos: Utokyo-sarulab system for voicemos challenge 2022

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.592094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.592094Z digest=sha256:02c2ce5e0132b885fbe7063b34d6ff83c9b165ed11bc0ce27fefdb13dc656890

Observation 6d485cd0-18c6-4f42-be48-006957a13a47 · outbound

This paper cites A toll quality 8 kb/s speech codec for the personal communications system (pcs).IEEE Transactions on Vehicular Technology, 43(3):808–816, 1994.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram A toll quality 8 kb/s speech codec for the personal communications system (pcs).IEEE Transactions on Vehicular Technology, 43(3):808–816, 1994

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.311192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.597654Z digest=sha256:f74aff760d764b2207b47b6faa11da06fa2dcfd6ce9bab6b72437d9a208f5cdd

Observation ad724549-5b05-493f-a20e-c0f221caf14f · outbound

This paper cites an unresolved cited work.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-12T18:47:52.294500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.603829Z digest=sha256:ddc2257979e9d72f50726b05867b680c151dc838a244fffab0c5fee144525742

Observation 223fd8ff-ccae-411d-ad3f-1723c05de0e4 · outbound

This paper cites Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.609553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.609553Z digest=sha256:15a1c42e61cd70b33bbae74ef590fb4764e18f3aa7162b9405382962e02599a1

Observation d25f7265-a625-4703-a6c3-af82bb254f83 · outbound

This paper cites Linear predictive coding systems.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Linear predictive coding systems

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.277542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.615331Z digest=sha256:24ac0716b1212d4432cac87cfa19afa0e49bbcd8f1ca69a0545ace9db2914d50

Observation 536dd418-5b3b-4fec-8080-3cb6b22ec255 · outbound

This paper cites High- quality, low-delay music coding in the opus codec.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram High- quality, low-delay music coding in the opus codec

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.261629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.621588Z digest=sha256:33b52e58298336173d1e937a8b276111562965771566d0988a8c29e6e45a96dc

Observation 455395df-d923-45ba-a8fc-d65ea5869f46 · outbound

This paper cites LPCNet: Improving neural speech synthesis through linear prediction.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram LPCNet: Improving neural speech synthesis through linear prediction

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.242851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.627258Z digest=sha256:a8b1f13bc11629c6136c1f231c34b62ff5d2bdc1afb079dcd2061336e16ba69f

Observation 5f834bf8-7453-4c3d-8c24-85198c114358 · outbound

This paper cites A review of vector quantization techniques.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram A review of vector quantization techniques

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.226450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.632471Z digest=sha256:38e0bdf1854b165cf4f22a7f556a9fc8a23b266f8eae64e9cef230ebc059dfeb

Observation 165e37aa-c693-4d38-a038-f46bb45f28ad · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.637690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.637690Z digest=sha256:6b27cc994f5d24fbcb1f459d6391b5973003c78370d766bbe71a88a789f1df41

Observation 6da65fae-c7cc-433d-92dd-8627fff1668a · outbound

This paper cites Neural source-filter-based wave- form model for statistical parametric speech synthesis.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Neural source-filter-based wave- form model for statistical parametric speech synthesis

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.209143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.644418Z digest=sha256:5729b3fcbbe62630bebc777a37e7451da4eb6954c187a730ef66b717ca5716b4

Observation 88c238f5-05fd-4d7b-b651-67814202c71e · outbound

This paper cites Tacotron: Towards end-to-end speech synthesis.Interspeech 2017, 2017.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Tacotron: Towards end-to-end speech synthesis.Interspeech 2017, 2017

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.191451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.650301Z digest=sha256:bd77d2d655ade6439ca67c6b946ddac7c0d426c885ee52fc6ba73e0e7e0e62bd

Observation 1bdeb114-39a5-4218-8e64-b04765d03778 · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoencoders.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Convnext v2: Co-designing and scaling convnets with masked autoencoders

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.172141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.655870Z digest=sha256:46dd62ff407a709562f29253f53cb4885bd76eb46aa9e82164ff8eda35ec482c

Observation dac23bcf-ef23-4e67-8590-0c402b8028b5 · outbound

This paper cites Audiodec: An open-source streaming high-fidelity neural audio codec.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Audiodec: An open-source streaming high-fidelity neural audio codec

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.154250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.662140Z digest=sha256:e8833e9e808c3df93478c90fdfb72e023a53055454bacd8fa291b1e388dfdca1

Observation 48aca32e-1094-499e-ad7a-dccea276c529 · outbound

This paper cites Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92).

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Cstr vctk corpus: English multi-speaker corpus for cstr voice cloning toolkit (version 0.92)

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:52.136766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.668031Z digest=sha256:e8aab803e374bc8031fb47c506d6852b4a055385447ae53a6a099146ff06047f

Observation eccdc361-0a0e-40e0-aae4-1271ce9acdc2 · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.673711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.673711Z digest=sha256:5a223c5a77d3cdbec7530858f32415655a1344648e5640655c53fdff44bfe8e0

Observation db58b2bc-4a36-4a1d-a59c-9283c5753b11 · outbound

This paper cites Source-filter hifi-gan: Fast and pitch controllable high-fidelity neural vocoder.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Source-filter hifi-gan: Fast and pitch controllable high-fidelity neural vocoder

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:51.989129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.678726Z digest=sha256:87aabea5b2156c914c85ee4c346cb6110fcc687c25d479ba45635edc5f6e02cf

Observation e86c4863-6802-4197-ad1c-3619e1441c71 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.IEEE/ACM Trans- actions on Audio, Speech, and Language Processing , 30:495–507, 2021.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Soundstream: An end-to-end neural audio codec.IEEE/ACM Trans- actions on Audio, Speech, and Language Processing , 30:495–507, 2021

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T18:47:51.970628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T18:47:51.683963Z digest=sha256:eb06f4e73993de227a4bb5e2a6559ab35430b2e8f003f8e094432fd9dd2f6695

Observation 0a7a1dab-4a31-4b46-9b73-f62714c0ae11 · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.689848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.689848Z digest=sha256:4ba245a71a79259986ebed9423ba957cf29d4196629fcdc989f7dee6e8bb5c10

Pith citing papers

No inbound Pith citation observations are available.