Pith. sign in

Paper Citation Record · LEDGER

DASB - Discrete Audio and Speech Benchmark

As of 6 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 9 inbound Pith citation observations for arXiv:2406.14294.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.14294 v4

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-24T00:26:57.419537Z

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T11:41:01.619120Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T06:36:02.212511Z

Reference resolution

75 of 75 outbound references displayed

  • verified exact15
  • verified fuzzy57
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 15c469c2-91df-4fbb-a195-4a58ce3ce92b · outbound

This paper cites Fundamentals of Speech Recognition.

DASB - Discrete Audio and Speech Benchmark Fundamentals of Speech Recognition

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.234372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:8809a62a638c35b2179155a868455c43e744473b9c97213ef153399a312010ff

Observation 13e199f1-2279-4d4b-91d2-1935c9b58de4 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

DASB - Discrete Audio and Speech Benchmark wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.100173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:59b896f30ba6f72e833bfd26c435368d1d424caac099706e7a11cf751637ee23

Observation a6cf519b-f990-451a-821d-99fb361eb3ea · outbound

This paper cites WavLM: Large-scale self-supervised pre-training for full stack speech processing.

DASB - Discrete Audio and Speech Benchmark WavLM: Large-scale self-supervised pre-training for full stack speech processing

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.132627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:b450315f34af61bfa81f339382a7f44bec8dbfdf6fe84da011fc95561855d9dd

Observation 2b909b64-7f53-439b-b256-e607c973aa57 · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units.

DASB - Discrete Audio and Speech Benchmark HuBERT: Self-supervised speech representation learning by masked prediction of hidden units

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.284575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:3d96e0a25040a69631d555488fe2927678f5e4175b3cef0e85a1e26a56b29820

Observation 0d5c2de1-6f10-4e47-8de4-43e537d6a67f · outbound

This paper cites Speech resynthesis from discrete disentangled self-supervised representations.

DASB - Discrete Audio and Speech Benchmark Speech resynthesis from discrete disentangled self-supervised representations

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.125574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:7428c8da9e370a0dd6665ab02e23c4e1eb876a896c51d6d594858b4a1a6988e5

Observation 8d849d4c-ac8a-458d-9d70-98bc83c7e5f2 · outbound

This paper cites Phonetic analysis of self-supervised representations of English speech.

DASB - Discrete Audio and Speech Benchmark Phonetic analysis of self-supervised representations of English speech

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.202045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:6a2e0837a98ac11c6ba14f885677f3b44a9854e8b2a8cd571d520adf5108cb77

Observation 70ce8680-2492-4182-b08a-062b7b6f52ee · outbound

This paper cites w2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training.

DASB - Discrete Audio and Speech Benchmark w2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre-training

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.180620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:82644307b76de5910e97fe6d6332a0bf3e8bfa314d06b3cf9d79ad497d29b3c1

Observation 8c993e73-6d70-4c7b-9087-54428fe589bd · outbound

This paper cites SpeechTokenizer: Unified speech tokenizer for speech large language models.

DASB - Discrete Audio and Speech Benchmark SpeechTokenizer: Unified speech tokenizer for speech large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.183800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:aee038ab7defebc3d9d431a828b2fea74d2ced62363e8c6f519fbe17f9ac6bad

Observation 9b041203-250c-4ff5-8dc3-10ede7c1c9bd · outbound

This paper cites FunCodec: A Fundamental, Reproducible and Integrable Open-source Toolkit for Neural Speech Codec.

DASB - Discrete Audio and Speech Benchmark FunCodec: A Fundamental, Reproducible and Integrable Open-source Toolkit for Neural Speech Codec

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:28:39.544200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:00c6b165814b9d42831343e798a83b2f92e7083d2af634944dbb8b01a57e123f

Observation 1703c728-23b5-4fcd-b342-a4b5afd9bd3a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

DASB - Discrete Audio and Speech Benchmark Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-24T00:28:39.538282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:0923e75007553a9144708db173cb20f742e1ee90d867b9786b735b79540e7359

Observation a86bf53d-c91e-4c10-8af2-ce7c2673e813 · outbound

This paper cites PaLM: scaling language modeling with pathways.

DASB - Discrete Audio and Speech Benchmark PaLM: scaling language modeling with pathways

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.194849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:f828f17ddb52acbec0f38203c7fb5593817fa90b9c3f4095b490d009a0d34db0

Observation 8a502682-a56e-454f-9839-88bb324ab983 · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding.

DASB - Discrete Audio and Speech Benchmark BERT: Pre-training of deep bidirectional transformers for language understanding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.173959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:d1fbaf6e87956e10c9ad6daefb6c5f1952b12612b252a23f332ea5c7a7f3e4f5

Observation 61f6c2c0-9f45-4c68-8b5d-71fa8bfa427a · outbound

This paper cites GPT understands, too.

DASB - Discrete Audio and Speech Benchmark GPT understands, too

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.241841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:11dff5f89903dd9e4768aac04dcab65ff9c92527207c41eccbcc76397eb7e9f2

Observation ca64e0c5-6ded-4842-8553-06062cf62582 · outbound

This paper cites Neural discrete representation learning.

DASB - Discrete Audio and Speech Benchmark Neural discrete representation learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.170126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:311e65e82128213f60a37b9b5304d752f92b9ea7400f499077e1acc155ba6017

Observation d49019ca-98ee-4907-bd2b-a437cbe31600 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

DASB - Discrete Audio and Speech Benchmark AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-24T00:28:39.570123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:c7f425246bf79d15429e7192ab0cfbf33310403309a50729f7bd4a42a8da4f31

Observation 1bbd4f02-3816-436a-89c4-61a94360561b · outbound

This paper cites LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT.

DASB - Discrete Audio and Speech Benchmark LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:28:39.499589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:de4efb34f2cd7c15cad913635a54c3b3e43beb2ef410a16a57d9c4f683bbdfc4

Observation c3e9db4e-2e37-4824-b626-ded906c5e4eb · outbound

This paper cites VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation.

DASB - Discrete Audio and Speech Benchmark VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:28:39.518355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:f4975164ae2630d1eca66069367b1078e5517c23c9d051ef21276ceadfe4f1fc

Observation fe4fb7ae-2b50-49c3-b335-f09f4c821bf1 · outbound

This paper cites SpeechX: Neural Codec Language Model as a Versatile Speech Transformer.

DASB - Discrete Audio and Speech Benchmark SpeechX: Neural Codec Language Model as a Versatile Speech Transformer

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:28:39.564746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:261a437bf585d3def47166d27593a0e643143fd0c9e09fbb5f770c5a9a8b5251

Observation 9336088c-dd6c-45f7-b706-52e8b9d35f21 · outbound

This paper cites MusicLM: Generating Music From Text.

DASB - Discrete Audio and Speech Benchmark MusicLM: Generating Music From Text

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-24T00:28:39.493046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:57913517053f5db320e48e7e7a668add1b714d01bc3d239f989af7447e580a07

Observation f34d4102-5c32-4525-bbba-77d67e58c8ee · outbound

This paper cites AudioGen: Textually guided audio generation.

DASB - Discrete Audio and Speech Benchmark AudioGen: Textually guided audio generation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.088490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:d345f262bc6968b92b263ff44252c242a28c297e1b2ebc28b4c8be85f59fd650

Observation 878435e1-6813-4057-ad83-7a42e56610e0 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

DASB - Discrete Audio and Speech Benchmark Gemini: A Family of Highly Capable Multimodal Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-24T00:28:39.549993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:0ca2cd4eb9ff9f7b2881c25a13e97a13417ba802cedb6ce166bb968a151d716c

Observation 0f4cde97-5bae-4f8d-8ab7-bf17bab37015 · outbound

This paper cites Goodfellow, Yoshua Bengio, and Aaron Courville.

DASB - Discrete Audio and Speech Benchmark Goodfellow, Yoshua Bengio, and Aaron Courville

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.288456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:945da42b6efca291a6234f8d985406ccff3565332345b64cd54911da78dc14db

Observation 8996c608-6bc4-4311-a090-db38c9c41f0e · outbound

This paper cites Codec-SUPERB: An In-Depth Analysis of Sound Codec Models.

DASB - Discrete Audio and Speech Benchmark Codec-SUPERB: An In-Depth Analysis of Sound Codec Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:28:39.559949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:456f1eccb255084ff1aa095db1341fc0a5a1986f312431292933e6524f2a6c76

Observation d223c66e-8160-421a-afde-b92740b529ff · outbound

This paper cites Puvvada, Nithin Rao Koluguri, Kunal Dhawan, Jagadeesh Balam, and Boris Ginsburg.

DASB - Discrete Audio and Speech Benchmark Puvvada, Nithin Rao Koluguri, Kunal Dhawan, Jagadeesh Balam, and Boris Ginsburg

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.238645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:86feff0cb8349a189879ad8cb2bf9a54c7a3334571f8750f05a5781a62a060bf

Observation 392c0a12-3309-446c-ba66-694ea6585373 · outbound

This paper cites SELM: Speech enhancement using discrete tokens and language models.

DASB - Discrete Audio and Speech Benchmark SELM: Speech enhancement using discrete tokens and language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.220672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:a7afd64febb6e9cce5930f10905ed417d2b15daa5ab320a28e2aea445b5f606f

Observation e79beb42-96e6-4af2-bf28-fac0fb5b4083 · outbound

This paper cites DUB: Discrete unit back-translation for speech translation.

DASB - Discrete Audio and Speech Benchmark DUB: Discrete unit back-translation for speech translation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.152518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:1a2d42f132a44e4d4160ef97b77d612249a9491bcd02b1ea0c8e1327638761c8

Observation e8ac3259-7c6d-4c37-a02d-019b4ed5beac · outbound

This paper cites Exploring speech recognition, translation, and understanding with discrete speech units: A comparative study.

DASB - Discrete Audio and Speech Benchmark Exploring speech recognition, translation, and understanding with discrete speech units: A comparative study

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.156112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:fc0806aa1e51833ed40ee8c146b9886a39b0c32f353febb7269c43749238689a

Observation ac437fc5-dfcf-4075-9840-1cfd8b0587b4 · outbound

This paper cites High fidelity neural audio compression.

DASB - Discrete Audio and Speech Benchmark High fidelity neural audio compression

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.217015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:6bcd88cfa9c233509391b174e4cb6a9d5100ec6ca970f2e6e2516f7205fad322

Observation e271ce55-f24b-41c9-8df4-e7079c5b4ff4 · outbound

This paper cites High-fidelity au- dio compression with improved RVQGAN.

DASB - Discrete Audio and Speech Benchmark High-fidelity au- dio compression with improved RVQGAN

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.209632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:f9896c278468787179fe6a289f1a84b4916a176234f3739c736bbd6cafe572c7

Observation 0062d389-d0ef-4c9a-ba48-659baf9bc9ad · outbound

This paper cites Speech self-supervised representation benchmarking: Are we doing it right? In Interspeech, pages 2873–2877.

DASB - Discrete Audio and Speech Benchmark Speech self-supervised representation benchmarking: Are we doing it right? In Interspeech, pages 2873–2877

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.159847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:e94c52fd9497e8f2cb826b2d5bd852f3509014d2d65c10937e3bc3cb10ee6326

Observation 03ff4e08-5398-4fa5-8d4f-0ee3c7af1055 · outbound

This paper cites SpeechBrain: A General-Purpose Speech Toolkit.

DASB - Discrete Audio and Speech Benchmark SpeechBrain: A General-Purpose Speech Toolkit

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:28:39.555237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:5e98180dc56a50b09bc74da9b073a81f848e226ab1c9930412175c67d6b8ed52

Observation d41dd303-ca44-4c34-9297-eb1c6b811810 · outbound

This paper cites Exploration of efficient end-to-end ASR using discretized input from self-supervised learning.

DASB - Discrete Audio and Speech Benchmark Exploration of efficient end-to-end ASR using discretized input from self-supervised learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.191380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:e5ec195406f93ae693a1d3750b30e441043622200a740b30cf0b3d3802670517

Observation 8890de73-44e7-4326-a5f1-27877e582cb6 · outbound

This paper cites Towards universal speech discrete tokens: A case study for ASR and TTS.

DASB - Discrete Audio and Speech Benchmark Towards universal speech discrete tokens: A case study for ASR and TTS

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.280909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:fd714020e4c9dfe6408cba126beca3c50db0e55b1251edabb280d52db008b610

Observation d61da613-17a8-4691-9ccd-770888baa129 · outbound

This paper cites TokenSplit: Using discrete speech representations for direct, refined, and transcript- conditioned speech separation and recognition.

DASB - Discrete Audio and Speech Benchmark TokenSplit: Using discrete speech representations for direct, refined, and transcript- conditioned speech separation and recognition

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.115703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:266dc687bedbc0fa20243f630279a1d2dfbcdd4835aa7b48459ccc473933d56e

Observation bba4bfbe-998e-45fc-b7ba-6ff79f8fe659 · outbound

This paper cites Evaluating text-to-speech synthesis from a large discrete token-based speech language model.

DASB - Discrete Audio and Speech Benchmark Evaluating text-to-speech synthesis from a large discrete token-based speech language model

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.135838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:77683277966ede05ebbd90ee1f3414536f4ad53807c307e66f2fe0373446fab3

Observation dfb0188f-cdd3-4cdc-9bcc-16a8bbe3cb0e · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

DASB - Discrete Audio and Speech Benchmark Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-24T00:28:39.486547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:f513e982e26edcf4cdffe85378a421129305357a2b5f047fc5fce86588ee7581

Observation d8d99b98-5bf0-4d82-94e7-b88a3502ee16 · outbound

This paper cites Speak, read and prompt: High-fidelity text-to-speech with minimal supervision.

DASB - Discrete Audio and Speech Benchmark Speak, read and prompt: High-fidelity text-to-speech with minimal supervision

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.119265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:912c32b6f4a4b0c2c40b0cf9534fbe7fb5396e74153c7f0914e646addda83656

Observation d6bbc6a9-16ca-4732-b4d1-9a1c5f56c263 · outbound

This paper cites How should we extract discrete audio tokens from self-supervised models?.

DASB - Discrete Audio and Speech Benchmark How should we extract discrete audio tokens from self-supervised models?

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.258681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:22d9e4fef263361f1d1d40fabaa60c738d3e821c443ba08f4d5c04ee92f7b022

Observation cc938772-cfdc-4df5-b454-5dcb5f94863d · outbound

This paper cites Lin, Andy T.

DASB - Discrete Audio and Speech Benchmark Lin, Andy T

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.231027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:647c5e49ecfb6bcefc073d665f82a1deca3f50cf4e8f9c38976fb3573f2e2fa1

Observation 1b64a4e4-8c7c-4cba-b1cf-629a9e320b0e · outbound

This paper cites Definition of the Opus audio codec.

DASB - Discrete Audio and Speech Benchmark Definition of the Opus audio codec

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.227762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:89118e16bba0a455d40ff58a2bd6bbfca638b9a7be5064720ef6e93d63ce0e37

Observation 1b3903e7-e8a8-4a85-af80-b8238cc91f44 · outbound

This paper cites Overview of the EVS codec architecture.

DASB - Discrete Audio and Speech Benchmark Overview of the EVS codec architecture

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.277353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:b6f1573f4a83dda7d42feb5f0b9db1659fed01cee59df73aac80e87e05c5d669

Observation efeb3452-7187-4465-9513-0c0801748e96 · outbound

This paper cites SoundStream: An end-to-end neural audio codec.

DASB - Discrete Audio and Speech Benchmark SoundStream: An end-to-end neural audio codec

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.107742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:3d41e77fa241011b2089cf8ce685d45d12b53691f2184a8338b3827b6f455af5

Observation 7c292212-8511-461c-9b90-db99a7ffb753 · outbound

This paper cites AudioLM: A language modeling approach to audio generation.

DASB - Discrete Audio and Speech Benchmark AudioLM: A language modeling approach to audio generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.112172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:f11977799362caf1a0df86f3b4e79a6b1e0b854ea8794b1a1d647c6c38c8394b

Observation 2db04592-9e87-449a-86a0-452f6c8ee9a9 · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

DASB - Discrete Audio and Speech Benchmark HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:28:39.528794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:61414fc86b85bf14c4776a2497e1d9a9653709cc2af8c0c11c0f319b12b814ca

Observation 9b449366-08cc-4d67-8ff4-aa9cd165b1b5 · outbound

This paper cites Free English and Czech telephone speech corpus shared under the CC-BY-SA 3.0 license.

DASB - Discrete Audio and Speech Benchmark Free English and Czech telephone speech corpus shared under the CC-BY-SA 3.0 license

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.129075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:f55a74a7ae7f1034661e875d2de33df7b117ae262d0e70c5aa850607f8a67268

Observation f7074f1a-e99f-4cb3-933a-0546eda1a35e · outbound

This paper cites ContextNet: Improving convolutional neural networks for automatic speech recognition with global context.

DASB - Discrete Audio and Speech Benchmark ContextNet: Improving convolutional neural networks for automatic speech recognition with global context

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.266311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:d3a0557c25a33427d768d01d014eda8ddfa52127d00140efce5c1740f03e3d54

Observation 1f521a2d-c557-4ab2-b634-123841be39db · outbound

This paper cites Common V oice: A massively-multilingual speech corpus.

DASB - Discrete Audio and Speech Benchmark Common V oice: A massively-multilingual speech corpus

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.262750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:8de6b346d9291f90e428443b07a60bc3383863a9b6260d6a59a85467080070eb

Observation 8729b9ce-8ecb-4755-a157-e84fb85e0029 · outbound

This paper cites V oxCeleb: A large-scale speaker identification dataset.

DASB - Discrete Audio and Speech Benchmark V oxCeleb: A large-scale speaker identification dataset

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.103616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:320aec093d7c1095aa1271168ec998ab621a7ef9f0a2ab2bfa3f58d0a5633293

Observation bdfd1d56-55f3-411e-8832-00afc782003f · outbound

This paper cites X-vectors: Robust DNN embeddings for speaker recognition.

DASB - Discrete Audio and Speech Benchmark X-vectors: Robust DNN embeddings for speaker recognition

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.122468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:4909a3a90c9e640ee8d72b0bc9d44918d5f55342a658c7d3f375a3cb2fa162ee

Observation 696670d3-987a-4bd7-9571-498ce8557aab · outbound

This paper cites Additive margin softmax for face verification.

DASB - Discrete Audio and Speech Benchmark Additive margin softmax for face verification

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.139489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:174ee94f92d667413c7da15cf8e27a29f33267db30b9c202748978cb2855cb78

Observation ac1e7933-f980-4694-8c89-a01c31ac370e · outbound

This paper cites ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification.

DASB - Discrete Audio and Speech Benchmark ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.251568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:cd0785a28d3236fbd794a85901dcbfa83a9123ccc54b80dae6abd72fc45ca82a

Observation 2753980d-eb41-47d6-ab9b-a298d4ddd2d1 · outbound

This paper cites IEMOCAP: Interactive emotional dyadic motion capture database.

DASB - Discrete Audio and Speech Benchmark IEMOCAP: Interactive emotional dyadic motion capture database

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.269814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:5a123ce354596911d4f85be8b29258b52dd4bffcc086839a70743e2f941ae931

Observation 3bbca81a-38a8-45c3-b377-7e9043e4e2cf · outbound

This paper cites SLURP: A spoken language understanding resource package.

DASB - Discrete Audio and Speech Benchmark SLURP: A spoken language understanding resource package

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.198549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:36e518276963658821b6e0e1de5a017493fd861c765c1f35fd4fc1529b70fc58

Observation beca2592-7b99-4ae7-bb51-172f3bb9e107 · outbound

This paper cites Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition.

DASB - Discrete Audio and Speech Benchmark Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-24T00:28:39.506695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:cf1e4f080c38412f401dfc17c4728d333ddb1b5589ec521728c5c69d94efd4fe

Observation 08fd36c0-2714-4c6d-b832-fb2fa54d4521 · outbound

This paper cites Investigating RNN-based speech enhancement methods for noise-robust text-to-speech.

DASB - Discrete Audio and Speech Benchmark Investigating RNN-based speech enhancement methods for noise-robust text-to-speech

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.273789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:514c274902f3bb92fecf932c1f7e6be3675d99e8f624c7009d4181db612038de

Observation 4cbc90bf-2d93-4b8d-a6ca-2b620b60cfe6 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition.

DASB - Discrete Audio and Speech Benchmark Conformer: Convolution-augmented transformer for speech recognition

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.187329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:f50b0706a1b56a47f25e4bab0856e4453dd60735875cec84d7318a00907df1cd

Observation 17039ad2-f2c4-4338-ac63-6725dc75cf54 · outbound

This paper cites DNSMOS P.835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors.

DASB - Discrete Audio and Speech Benchmark DNSMOS P.835: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.245085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:bb45983468eac98953ab584f9613e9f2f5ba538ffa3c263cf0be6d25863a4900

Observation 728b533b-f475-46da-b564-61f7f90b1b76 · outbound

This paper cites Sequential multi-frame neural beamforming for speech separation and enhancement.

DASB - Discrete Audio and Speech Benchmark Sequential multi-frame neural beamforming for speech separation and enhancement

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.177334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:0a061faa4feef34b88f9f1724def48f76185c321998ac6ede7ca2c4e7f7f75c9

Observation e3f41e3d-44d2-462e-8ef7-e2a26197d103 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

DASB - Discrete Audio and Speech Benchmark Robust Speech Recognition via Large-Scale Weak Supervision

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-24T00:28:39.523430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:ddb7d94bc2c41b3c79456a1c4eae0b8342183f984481cafe6098fda10eabcff1

Observation 5332612a-c249-475b-b45f-46194f927f85 · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

DASB - Discrete Audio and Speech Benchmark LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:28:39.512180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:1fc397397bf27b7c4e5240e2f770636894732fe9d6d9f799adab99f6af7fd768

Observation 08c5f7d4-328e-4f45-b086-74f70d493221 · outbound

This paper cites Kolbæk, D.

DASB - Discrete Audio and Speech Benchmark Kolbæk, D

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.162502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:b95f4446e506050a75c80e567c038e84f3e8b985072cca72a0c1ac1739cfbe6c

Observation 4e96f246-c73d-4e38-969f-1db824fa6678 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

DASB - Discrete Audio and Speech Benchmark Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.255111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:e4d9a32fcf39592c85d3fedef279c984cc65ddd380a2b057b08fd9d6529c5b38

Observation e3178b6d-9be5-40bb-baec-8be39f299650 · outbound

This paper cites The LJ speech dataset.

DASB - Discrete Audio and Speech Benchmark The LJ speech dataset

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.223867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:460c3499b0b5fbdd16c277027d436a3b3c6c33dd81046c7317df9640d8417f41

Observation 5421ef39-58ad-42b0-a393-29c4bf81ab23 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for V oiceMOS Challenge 2022.

DASB - Discrete Audio and Speech Benchmark UTMOS: UTokyo-SaruLab System for V oiceMOS Challenge 2022

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.096920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:a82a50661175668a597858ba494dd104ae88e158713f6253ba5ff80e917d9450

Observation 78d5c962-7ff3-423b-926c-c96e48e8c04e · outbound

This paper cites A comparison of discrete and soft speech units for improved voice conversion.

DASB - Discrete Audio and Speech Benchmark A comparison of discrete and soft speech units for improved voice conversion

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.084365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:47e9a601533f49576350f5ca57a3092e7909198d83322f81eef304c21ca2aa58

Observation fcaae5e3-9967-45b6-8701-063e49613ee3 · outbound

This paper cites Gebru, Dejan Markovi ´c, and Alexander Richard.

DASB - Discrete Audio and Speech Benchmark Gebru, Dejan Markovi ´c, and Alexander Richard

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.248489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:2945ad39d064d438c82d0f8a390a26879ccad7e0661b6d224c9427fb50e9a8d7

Observation fcb2830f-346e-4df1-b33e-94cb24161d44 · outbound

This paper cites ICASSP 2023 deep noise suppression challenge.

DASB - Discrete Audio and Speech Benchmark ICASSP 2023 deep noise suppression challenge

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.146255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:3de3fd1b29f022b936de4329671ff600b7c81e513d360df10332658ead7641f4

Observation 0c858094-83eb-40f1-84a6-7d5900267575 · outbound

This paper cites Audio Set: An ontology and human-labeled dataset for audio events.

DASB - Discrete Audio and Speech Benchmark Audio Set: An ontology and human-labeled dataset for audio events

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.213448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:ae3ebc5298885a2dc08f608bf6ccf2071503d014dd43b6bbeedd849709907a70

Observation 8edd5644-5573-4186-895d-bc025d0813c5 · outbound

This paper cites FSD50K: an open dataset of human-labeled sound events.

DASB - Discrete Audio and Speech Benchmark FSD50K: an open dataset of human-labeled sound events

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.093017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:cdda79f61fc1642941ccb7de214844c74d861bc7a0a1e31f6d692975f320b58c

Observation 8a4da591-7b4b-4eb8-8f98-788a4136c2bb · outbound

This paper cites The MTG-Jamendo dataset for automatic music tagging.

DASB - Discrete Audio and Speech Benchmark The MTG-Jamendo dataset for automatic music tagging

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.142921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:78b2f9a2b4913eb172d51040fb78b93709fcc1ac6aa6f56314546140be541136

Observation 3a7d9e73-d2ad-4b91-8772-94b8c34cd542 · outbound

This paper cites an unresolved cited work.

DASB - Discrete Audio and Speech Benchmark Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-05-24T00:28:40.149411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:18b05abc91d461c15d4a9aebe556ca5bf2918e3981c75f5630fa1f9c54d126a6

Observation 18b2400e-55ef-4dcb-9c80-e27031a228f9 · outbound

This paper cites CSTR VCTK Corpus: English multi- speaker corpus for CSTR voice cloning toolkit (version 0.92).

DASB - Discrete Audio and Speech Benchmark CSTR VCTK Corpus: English multi- speaker corpus for CSTR voice cloning toolkit (version 0.92)

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.205787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:5b62250243500e9a0a2ed4a7b4a926f48134459894b597da6943787f4543e73f

Observation 10e9b2d1-babb-4432-9c04-e17a3ab65b7f · outbound

This paper cites I., and Bittner, R.

DASB - Discrete Audio and Speech Benchmark I., and Bittner, R

Reference 74

Resolution
metadata mismatch
doi, observed 2026-05-24T00:28:39.268137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:bbcfb12cb83eaf0b39ff621a1ff40862aee02ff91b2922cb8a66b08c7be96474

Observation 6656e935-0029-4694-a116-4d5582943028 · outbound

This paper cites Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, Rif A.

DASB - Discrete Audio and Speech Benchmark Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen, Yu Zhang, Yuxuan Wang, Rj Skerrv-Ryan, Rif A

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T00:28:40.166268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:3fde939b486204f536028f93aa32858cc1be26819c954a2522709cad601350e8

Observation 29e9a675-7200-468e-8b7a-93fd7d99f902 · outbound

This paper cites Not Converged.

DASB - Discrete Audio and Speech Benchmark Not Converged

Reference 76

Resolution
malformed identifier
arxiv_id, observed 2026-05-24T00:28:39.533841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:2f1ee14cd5f4048683d2cff45795545fe27580e8a5dbf264d5a86c96cff3496d

Pith citing papers

Observation 7a10a063-3204-4ff9-a0a3-3b95c7d7218d · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey DASB - Discrete Audio and Speech Benchmark

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-22T20:45:08.140406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:90a18b90c9a7b83a884a1d1b9d52cfd3ad3506af18117448cc48e98e4792c0a8

Observation 9f7dcdd6-2138-4130-9629-02d504a92f00 · inbound

AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation cites this paper.

AudioCodecBench: A Comprehensive Benchmark for Audio Codec Evaluation DASB - Discrete Audio and Speech Benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T11:41:01.619120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:41:01.619120Z digest=sha256:a4aa86b8bd7b2f68e2be9c167d451b4c7c49f5b65a53bca4b9a0bd54c8417f60

Observation cb31279b-701c-4689-9928-3e354ea8e814 · inbound

Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission cites this paper.

Analysis of Speaker Verification Performance Trade-offs with Neural Audio Codec Transmission DASB - Discrete Audio and Speech Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T11:29:02.089391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:29:02.089391Z digest=sha256:3b04cfce724051bb4417f69905680db3d61743effbe3bef35ceaa2cab28c786f

Observation 9127ad96-fd4a-4e9c-b4af-12c39d7984c8 · inbound

An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training cites this paper.

An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training DASB - Discrete Audio and Speech Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:53:40.339604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:53:40.339604Z digest=sha256:1a610f25f7288b9ff3b20f1225665990b8ce90cbf1a008037873430781407d22

Observation 2edffd4a-e698-4036-bcf2-a75547165f30 · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs DASB - Discrete Audio and Speech Benchmark

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:01:24.351659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:0bfb6b3eb156751e424f70834f05396301d95fe5a34fcebca652e0c842286997

Observation c1637969-f234-4c78-9fe8-b2521552ec14 · inbound

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception cites this paper.

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception DASB - Discrete Audio and Speech Benchmark

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T22:32:44.576420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T22:26:20.100595Z digest=sha256:5c556a8df18f73995d991736e92bd227ee2a840c0382a2f7354bb426fb5d7f73

Observation f81e9906-fbc7-48d2-b33d-e21087082546 · inbound

HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models cites this paper.

HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models DASB - Discrete Audio and Speech Benchmark

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T01:02:56.239274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T00:56:43.991936Z digest=sha256:7f5d595af085ed4f03d1a3190cfcb25ea767ef102503174fd61f0039d4e4bac0

Observation 7c82dd32-a45d-4026-9a6d-5528445a6d95 · inbound

Text-Independent Speaker Verification Using Discrete Audio Tokens cites this paper.

Text-Independent Speaker Verification Using Discrete Audio Tokens DASB - Discrete Audio and Speech Benchmark

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-09T06:36:02.215111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T06:26:45.962895Z digest=sha256:7b5652e1e9039b0926bff113dbbbda2dae575ba20693665616189e57fadeacfe

Observation 16c8ef31-d6ae-4f02-87cf-2d42377653f3 · inbound

Text-Independent Speaker Verification Using Discrete Audio Tokens cites this paper.

Text-Independent Speaker Verification Using Discrete Audio Tokens DASB - Discrete Audio and Speech Benchmark

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T15:43:46.809326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:43:46.809326Z digest=sha256:d20447219cf6fc75447396d738572540332cda4b75b2d2f5ef475ca25b967494