Pith. sign in

Paper Citation Record · LEDGER

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder

As of 11 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2501.05332.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05332 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:18:15.195399Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78316f17-55a2-4429-872a-b6473ab948d5 · outbound

This paper cites an unresolved cited work.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T21:18:16.056272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.921168Z digest=sha256:c2fc176c0250be7d698765435e2cd6e6a981b17b1f4227bd2cb168f1ac405bdd

Observation cf939511-44de-4644-af16-085bce42c0ac · outbound

This paper cites Speech analysis/synthesis based on a sinusoidal representation,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Speech analysis/synthesis based on a sinusoidal representation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:16.039414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.927861Z digest=sha256:1fa9a46b211fcd2546ce0d29bf8bb85e4b65f7012490454899fde5a24e13ec42

Observation c97c97b4-7d2d-4cca-8193-6c69c27c2290 · outbound

This paper cites Speech analysis/synthesis and modifica- tion using an analysis-by-synthesis/overlap-add sinusoidal model,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Speech analysis/synthesis and modifica- tion using an analysis-by-synthesis/overlap-add sinusoidal model,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:16.021198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.933873Z digest=sha256:61dc75743d0357ce7f200d73576cf788d2c93bbaf77504ef307bf2dc0d71742d

Observation 8a77f054-5119-4a95-a723-b81134781679 · outbound

This paper cites Spectral modeling synthesis: A sound analy- sis/synthesis system based on a deterministic plus stochastic decomposi- tion,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Spectral modeling synthesis: A sound analy- sis/synthesis system based on a deterministic plus stochastic decomposi- tion,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:16.001357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.939389Z digest=sha256:de9cbfa7bc4ca1d3ef68c14464da0ee4b078530426e6864395aa3a5ca5762fb5

Observation ab84d7e6-ad06-4345-87f5-2bb589433605 · outbound

This paper cites HNS: Speech modification based on a harmonic + noise model,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder HNS: Speech modification based on a harmonic + noise model,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.977931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.945301Z digest=sha256:493ff5db6d2a64d3b94f83285bad6320f2da99b1b8112e4821785119bb3b3283

Observation 188a3b83-9118-433d-a814-fc430dc3649b · outbound

This paper cites Analysis/synthesis and modification of the speech aperiodic component,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Analysis/synthesis and modification of the speech aperiodic component,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.957651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.950661Z digest=sha256:b4f058c6999385a058dbfe0f611b99966fb6b0f7d26c21bff1c6ae14b1146110

Observation 9455ab30-511c-45e4-9d69-20246d1ccfd7 · outbound

This paper cites Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Pitch-synchronous waveform processing techniques for text-to-speech synthesis using diphones,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.937209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.956552Z digest=sha256:a74483c31a2aa0395f226b77b8c9cdde56f98129f9246d1029732bd2d2ab270c

Observation e7f06493-4983-4df9-bd53-89ec86b944a3 · outbound

This paper cites Improved phase vocoder time-scale modi- fication of audio,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Improved phase vocoder time-scale modi- fication of audio,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.911352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.961508Z digest=sha256:f5b8817129986ad5a62fed346fc3c046bc7f93d761ab5a80d9e3febdc88a8072

Observation e74256c8-99bb-49da-89e2-0240c9d9bdda · outbound

This paper cites STRAIGHT, exploitation of the other aspect of VOCODER: Perceptually isomorphic decomposition of speech sounds,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder STRAIGHT, exploitation of the other aspect of VOCODER: Perceptually isomorphic decomposition of speech sounds,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.893689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.966599Z digest=sha256:016ca4ce2767cb9901a5cb33d157cfd0767f149ac7cfb921ecaaf9e58155e8d9

Observation 09759bbd-2886-4718-be68-de420a8b5dc9 · outbound

This paper cites World: a vocoder-based high- quality speech synthesis system for real-time applications,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder World: a vocoder-based high- quality speech synthesis system for real-time applications,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.874649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.971913Z digest=sha256:020519f252bb237b878467959ce7f18912cde6f7e938f4669d18c4bddcfe29a9

Observation ef9f6a6b-fccc-4a3c-a941-b8d69dc941d3 · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Soundstream: An end-to-end neural audio codec,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.854015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.976248Z digest=sha256:4c8c9e71dbe3da401b394edb727f9b832ad49d829d59fc2748f3b422ecbcfb0d

Observation cbeca490-e130-432e-89c3-f35dc3aa61a0 · outbound

This paper cites High fidelity neural audio compression,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder High fidelity neural audio compression,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:14.980882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:14.980882Z digest=sha256:0a671c06a43adba61d43237af3491ebe7b584af2476e6ab89124c27defa1c7d4

Observation caa1de04-6e11-47e2-a2ca-af8b7e9b465e · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.822520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.985768Z digest=sha256:99c1c7e01563aab7689548db8d40b20d5358bf5cafee8c1687f018f7fd8c2b06

Observation c9521e5c-be4e-42b5-b8a2-e72bb11f31ff · outbound

This paper cites Speech Resynthesis from Discrete Disentangled Self-Supervised Representations.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:14.990854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:14.990854Z digest=sha256:5fba0c4f1552d0ef5b90917f2dd6626928e0faf1dc464c468dbd36546991ea43

Observation 87ea4ccf-987b-4777-9fd3-b96796e9e6a6 · outbound

This paper cites Masked autoencoders are scalable vision learners,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Masked autoencoders are scalable vision learners,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.803457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:14.997002Z digest=sha256:83cd9e4573b48d60a785aa94bf278eb423bb8227009bfae3c85422cf113559c7

Observation cfa6c315-f0c5-4f3c-ad6b-037b4117194d · outbound

This paper cites Hifi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Hifi-GAN: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.786693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.011918Z digest=sha256:78c9ce58c8161fc25454bacabe1f42183d7c9d0a0c6f5e9e3817f251fda677b6

Observation 764b5b8e-019d-448b-b4a8-547843a84ddd · outbound

This paper cites CREPE: A convolutional representation for pitch estimation,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder CREPE: A convolutional representation for pitch estimation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.763028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.020133Z digest=sha256:eb0c6fbd4038fa76e32bf67aedbc0bbb2d21b6820135f6ae397b6defc1a18e83

Observation f2252bfe-383e-4dcd-8ce8-a514fd94867f · outbound

This paper cites ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.025829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.025829Z digest=sha256:1f23c298ac1b28a51cf9a633078060491c1df9f8e9b87ea203dd792f2bf5ff37

Observation be57210f-98ea-44cc-8da8-2b47308fb7bc · outbound

This paper cites Brouhaha: Multi- task training for voice activity detection, speech-to-noise ratio, and c50 room acoustics estimation,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Brouhaha: Multi- task training for voice activity detection, speech-to-noise ratio, and c50 room acoustics estimation,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.747351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.032234Z digest=sha256:8462047199c1648e9994a2f307228514609ea7b9b69401b642815cc76b5d9ab7

Observation cd3d3ccd-d69f-42f2-86e1-a075dd9b3454 · outbound

This paper cites Neural discrete representation learning,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Neural discrete representation learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.731382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.039976Z digest=sha256:588c8304307da6644764ac2e60bba2644cbfaceb922b0c569baba789681e35dc

Observation 0131dad0-d0cc-41a3-9ce5-e2bf3e9f6390 · outbound

This paper cites A vector quantized masked autoencoder for audiovisual speech emotion recognition.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder A vector quantized masked autoencoder for audiovisual speech emotion recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.048903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.048903Z digest=sha256:79e3610a3caa002509bcaddfcf832f5630edb641f659415bdc583fb4f4c6013d

Observation 5b12b587-a5e6-4c31-a7e9-00eead8f41b0 · outbound

This paper cites A comparison of discrete and soft speech units for improved voice conversion,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder A comparison of discrete and soft speech units for improved voice conversion,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.715030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.055931Z digest=sha256:bc8f09196d0be2d8c43e1a8878ff3c3f02a5234a8e6b81359038ed62c432737f

Observation 9c9d8261-2a90-4ef7-9df1-48e127b2fe7b · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.697722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.061930Z digest=sha256:aa18b14cbd83d00785b4788fa2e9807ad015a4cfcd757aa18d68287d98315053

Observation 3c0d0dc3-acb0-4387-a660-9bafcd021556 · outbound

This paper cites Attention is all you need,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Attention is all you need,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.682108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.069225Z digest=sha256:5474e141d7f85680a959e29d39ce4495b1c090539077306a8bfa967c2bad2f85

Observation ead23dc7-9872-488e-bf7f-f032c23489c4 · outbound

This paper cites DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.078041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.078041Z digest=sha256:6e910cbd8033271dfac202110e1865279011b3dd27918275539b6ee85fdadad4

Observation 512054d5-6ed5-4586-83dc-012f29817b1a · outbound

This paper cites Phase- aware speech enhancement with deep complex U-net,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Phase- aware speech enhancement with deep complex U-net,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.666451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.084578Z digest=sha256:29a3504d641dc386fd72c0d412db046ceb6daf3c4ad645f4965a148ea6ace0da

Observation 0de9e15e-57ae-4738-aec3-aba81e331c3c · outbound

This paper cites Dual-Path Transformer Network: Direct Context-Aware Modeling for End-to-End Monaural Speech Separation.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Dual-Path Transformer Network: Direct Context-Aware Modeling for End-to-End Monaural Speech Separation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.089856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.089856Z digest=sha256:a352cb903ec996545ed049fbe74d913ab9cc19b486a6b2127225616b052b2602

Observation c94154bf-a919-4a22-848b-827755ed68b3 · outbound

This paper cites Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.649768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.095540Z digest=sha256:81f245299d30ffdb7814f3a3ab66e9923de1782b056592145e956f7cd78cfa4e

Observation a1c9c67c-0775-47b4-b998-8cc8c4419e08 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Librispeech: an asr corpus based on public domain audio books,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.632143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.101155Z digest=sha256:b41ce8e45c05bc099620bde175b25b279bb8deb46f9424b7a6adc29fa52fde72

Observation 1b54d229-4a4f-46f9-8fd5-7a8766f7cacc · outbound

This paper cites DEMAND: a collection of multi- channel recordings of acoustic noise in diverse environments,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder DEMAND: a collection of multi- channel recordings of acoustic noise in diverse environments,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.616365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.106247Z digest=sha256:610beacaef6fe90fa5872f6ee3efe6da510d2b5e3ad7661e0a00d12796ae8cfd

Observation fecfb541-5dc9-4496-b178-08330e6a079a · outbound

This paper cites Pedalboard,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Pedalboard,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.110816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.110816Z digest=sha256:c675c1a471c9160146eaa6aa5c26ff40cfa24e6867304f39478dd822fd2443be

Observation 2f270ae9-7702-4761-bae5-017a66649232 · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.115417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.115417Z digest=sha256:3b94e6f72fb9e6c1a60b824584bbac00398d63659e27448301bac7ee5aa48ff5

Observation a8e88527-72e1-4c09-890a-5cc3a0765c84 · outbound

This paper cites Wham!: Extending speech separation to noisy environments,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Wham!: Extending speech separation to noisy environments,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.599067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.120486Z digest=sha256:1232e92355f64a73f40718f3f87537b77468cf580e8dd62f699466a142f1ade5

Observation d47231e4-35b9-4781-a762-26a42955d8a6 · outbound

This paper cites Decoupled Weight Decay Regularization.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Decoupled Weight Decay Regularization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.129971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.129971Z digest=sha256:93e1b118b5eb2268ae602a1629b23df00bdaf27f21d49e93a5618a26b85632fa

Observation 2f8f3491-dbb2-4707-a545-67bf629a53cd · outbound

This paper cites SDR–half-baked or well done?.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder SDR–half-baked or well done?

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.581834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.136261Z digest=sha256:4834da876b790b4cc0a69e887b53cc9cf23b57201008f22ed52c63ee617636b3

Observation 3b48262a-ffeb-4c02-bfa6-c10ab9d26b63 · outbound

This paper cites Perceptual evaluation of speech quality (PESQ) – a new method for speech quality assessment of telephone networks and codecs,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Perceptual evaluation of speech quality (PESQ) – a new method for speech quality assessment of telephone networks and codecs,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.566373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.142181Z digest=sha256:d53789af07c8044c0598fe95641721360a0c34b90115e29114d4e2157ba0c42e

Observation 8b20d736-d8d6-40eb-b500-868462a92682 · outbound

This paper cites An algorithm for intelligibility prediction of time–frequency weighted noisy speech,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder An algorithm for intelligibility prediction of time–frequency weighted noisy speech,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.550740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.147179Z digest=sha256:3d193cad50addd42f0a1530dcec287d52a7510e2769faa093ee054e2bab387d5

Observation d7e7dd40-04d7-4d69-a8f0-bca3b4530130 · outbound

This paper cites Generalized end-to-end loss for speaker verification,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Generalized end-to-end loss for speaker verification,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.534553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.152200Z digest=sha256:b75e440d903ec51e91ca2bb67ec14e72f8508042e94a2b38f6ce18f101e1595f

Observation 45192b41-6934-44e2-8e63-2e92347c00bb · outbound

This paper cites Speech quality assessment through MOS using non-matching references,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Speech quality assessment through MOS using non-matching references,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.517122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.157220Z digest=sha256:7845471d1c3acca4d0c508986ae88587ed23bf8c73c0a082137f15d4145e8d6c

Observation 1f76d720-2f26-40d7-9d5d-f9e853a80288 · outbound

This paper cites Torchaudio-Squim: Reference-less speech quality and intelligibility measures in Torchaudio,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Torchaudio-Squim: Reference-less speech quality and intelligibility measures in Torchaudio,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.500396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.162432Z digest=sha256:6ca6621fe3b016500861dd5a09d8393b813bb4de98b2612c4ef1c0dd06bc45b9

Observation ce955239-e364-4fda-991a-f0bf358e0c7a · outbound

This paper cites DNSMOS P.835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder DNSMOS P.835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.482873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.167258Z digest=sha256:583484e15af47d27a203f677496e03444802088264d1c909cd871d267a3b2269

Observation 7ed81e39-7248-4c77-a725-851eead3c682 · outbound

This paper cites A pitch tracking corpus with evaluation on multipitch tracking scenario,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder A pitch tracking corpus with evaluation on multipitch tracking scenario,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.465591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.172809Z digest=sha256:769fedf833854c41f09f7fecc59b86711f9434f9b206eef5f6a114c805594f2e

Observation 5b793503-8cbb-4b0c-8407-9945181c9317 · outbound

This paper cites pYIN: A fundamental frequency estimator using probabilistic threshold distributions,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder pYIN: A fundamental frequency estimator using probabilistic threshold distributions,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.445626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.178371Z digest=sha256:0439fdf8a00976d427c2db46e7b2c8a0f130701d33515efa8bd70e3b16760427

Observation e31f8484-8a53-46c8-b473-90725fca49ec · outbound

This paper cites A sawtooth waveform inspired pitch estimator for speech and music,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder A sawtooth waveform inspired pitch estimator for speech and music,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:15.184183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:15.184183Z digest=sha256:7e4e4f2eceff1ca01577884e35f04aeff7abd1bd30d4efb6e77ae95b9d8ff923

Observation 3016e0d9-78bb-43c2-b5bf-7450912cd148 · outbound

This paper cites Revise: Self- supervised speech resynthesis with visual input for universal and generalized speech regeneration,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Revise: Self- supervised speech resynthesis with visual input for universal and generalized speech regeneration,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.412512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.190587Z digest=sha256:7c1a60f0e067f3f212b732412ee3136a57f75690d949d720ea2dad96b3a8066d

Observation 6cda7d60-be83-49b5-84b9-70df58dceec9 · outbound

This paper cites Generating diverse high- fidelity images with VQ-V AE-2,.

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder Generating diverse high- fidelity images with VQ-V AE-2,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:18:15.395111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T21:18:15.195399Z digest=sha256:bedb12ef831b1117523e244e2cc58d4b4372661465cf16b4fbe20c3dbf147ed5

Pith citing papers

No inbound Pith citation observations are available.