Pith. sign in

Paper Citation Record · LEDGER

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music

As of 13 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 3 inbound Pith citation observations for arXiv:2412.11449.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11449 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:00:02.820379Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:00:02.477825Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:29:04.296162Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact3
  • verified fuzzy16
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 362bc031-fc62-45fd-8ce4-a426307fae6b · outbound

This paper cites an unresolved cited work.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:00:04.076314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.467848Z digest=sha256:cf028a67e69f69270c9cd08d241941cd4dd188d999aef70f196ecc9b9d544dfb

Observation bc0db306-bc06-4066-b8ba-52afd7addff8 · outbound

This paper cites For the case of speech, we use the LibriSpeech TTS dataset [27].

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music For the case of speech, we use the LibriSpeech TTS dataset [27]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:04.040769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.488536Z digest=sha256:89a1a34432459531c313759801e4688574d2363cec02926c56d628b3d234e0b5

Observation 2e5c4b4d-1b49-4c65-9fac-592d68c1abba · outbound

This paper cites an unresolved cited work.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:00:04.011740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.494724Z digest=sha256:aab0a089ca826413f6029c76e0ad6d9e113da0c4dc40277c90eed07c7f174eed

Observation 7729d3d9-6cc1-4807-8946-76d1527f2000 · outbound

This paper cites We are only interested in quantifying the likelihood scores for the gener- ative architecture pre-training.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music We are only interested in quantifying the likelihood scores for the gener- ative architecture pre-training

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.932322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.514037Z digest=sha256:4ba869836614049c8f49f22f92bb67af489298f25449f13fcff97c143f5ff24b

Observation 82ae8649-7b3c-48fc-9f31-ec5132a82eae · outbound

This paper cites The proposed architecture outperforms a purely token-based model by combining continuous audio and dis- crete token-based representations.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music The proposed architecture outperforms a purely token-based model by combining continuous audio and dis- crete token-based representations

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.900149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.521600Z digest=sha256:2466e029388a78eb0eff309bd42107778c916ff5f08d2b96ea4c19791fb6f770

Observation e785a6ed-4fb8-4a30-a4ae-af7a8410fb0d · outbound

This paper cites We thank both Google and Stanford HAI for this initiative.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music We thank both Google and Stanford HAI for this initiative

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.876394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.529451Z digest=sha256:9a014f506c6289453545f3b83f5c13f86829dbf6975806f894e7bd1f0efb7534

Observation f5cda316-aa4f-4aa5-852c-433e36aeb251 · outbound

This paper cites Audiolm: a language modeling approach to audio generation,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Audiolm: a language modeling approach to audio generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.749998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.605338Z digest=sha256:2fa8b85fa0adbdfa848fb1d7e0693077d9ccb2a2109608ead18eb1a16d9c8599

Observation f665c2dd-1a20-43f4-8462-a91236f5a276 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.613792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.613792Z digest=sha256:429625ee91b6e8e0885f7c404b0c2686ea0d97109114838d9f99e68fb149184c

Observation 7658836d-6cc0-4610-a070-d02b1bfd9d85 · outbound

This paper cites Attention is all you need,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Attention is all you need,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.845190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.543785Z digest=sha256:a17b01cc253a8f4802461ade506d328f75c10f589a65abf1823373bae8200c8e

Observation 63d715bd-9b42-419d-b9a0-c1defdcb6030 · outbound

This paper cites Language Models are Few-Shot Learners.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Language Models are Few-Shot Learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.561304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.561304Z digest=sha256:25babc54dc32edc7bdc312ae43b788137ffa375dca6e7baac5a738243a7f67ce

Observation b6b9e469-c7c9-4de3-8093-7581663fd5c5 · outbound

This paper cites Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.477825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.477825Z digest=sha256:6a6c89fab9784ddc3ab7a127108e9d7daf3f3f854437f7333e0ed1343f17ec53

Observation 41a5132d-32c1-4dce-aa0c-dbcdc6375438 · outbound

This paper cites A generative model for raw audio using transformer architectures,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music A generative model for raw audio using transformer architectures,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.814812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.569787Z digest=sha256:a36ccdcdd91b23c18f926266b5f4e20e77998c96cbafaa0db2563e20143f9029

Observation 1948d389-e4ef-4ffa-b2d3-bdf9fdae0a7d · outbound

This paper cites A Language Model With Million Context Length For Raw Audio.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music A Language Model With Million Context Length For Raw Audio

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:00:03.384865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.579250Z digest=sha256:c64805a5ebfdc813c6b788264e976b492b59af207db02017cd3f541d6ee1b4a5

Observation 852147a6-0030-4b57-934d-0f67c8d6511e · outbound

This paper cites Music transformer: Generating mu- sic with long-term structure,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Music transformer: Generating mu- sic with long-term structure,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.784848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.587047Z digest=sha256:f9510eed2f8e282a378cfeeab23e971eca4744c04f350b89a26410b8890c25d2

Observation 9cea3149-d154-4584-a557-582ba3c8dfde · outbound

This paper cites A Framework for Generative and Contrastive Learning of Audio Representations.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music A Framework for Generative and Contrastive Learning of Audio Representations

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:00:03.352827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.593871Z digest=sha256:0efeeb2d96adb356f9bd669d374cc335e781db76b6f359d4a79bd5d345a323a7

Observation d2bfc67e-447f-46ff-a2fb-684b61a9c752 · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music WaveNet: A Generative Model for Raw Audio

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.704954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.704954Z digest=sha256:915f0e218785d903b49b688566ea18bc1faece05a0d3547489b6d22566ba7de2

Observation 58f88346-8bc6-4d2c-94ee-27060dc810a6 · outbound

This paper cites Audio Transformers.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Audio Transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.621252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.621252Z digest=sha256:12917660e6996ad931fda6811218537b59c6d043067f42c09825af86cd4cf3ec

Observation 8384468e-a393-4ee6-aca8-843f12b7a0ba · outbound

This paper cites Psla: Improving audio tagging with pretraining, sampling, labeling, and aggregation,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Psla: Improving audio tagging with pretraining, sampling, labeling, and aggregation,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.721475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.629242Z digest=sha256:22ebfbeff07d5744de6e6c3b780707917386dfec3687f39f1ff78bcd1b5764d3

Observation bc88afda-ab7d-4059-81e1-54b6faafef48 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.692846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.637820Z digest=sha256:03555c0263cf97684df3180b47bfcebfc84bcaf2ce70d4f9d179e403285dbc80

Observation 94d345ac-4436-4f86-95d8-d32f6e75e212 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Gemini: A Family of Highly Capable Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.646427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.646427Z digest=sha256:e6ea0a585304fca51d217a0898fd07c615590972b69254d719a5832eb303893e

Observation 94c9d3b6-9c77-4c60-988c-e384a9c0773a · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.667660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.667660Z digest=sha256:da063d97fbc81c5ce7bd5357ff26e8fdd33da2191daf0b3899158a1195080cf0

Observation 8544ca89-6801-4570-89f0-1bfd1474ca8c · outbound

This paper cites Neural Discrete Representation Learning.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Neural Discrete Representation Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.684978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.684978Z digest=sha256:f97a42a0dbdc017436079b201cb0caf471806c5e6ca72e834fc6f224d97a9228

Observation 79d1a1d4-5e7f-4036-af3a-d60c1ad548f5 · outbound

This paper cites an unresolved cited work.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:00:03.964088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.502571Z digest=sha256:184acdf4364549436a1f1736d77cddf0ad6c97ce2d16b5f87cc45c156f985155

Observation 8aae6c97-b2d8-4cbd-a690-ec6025554102 · outbound

This paper cites Jukebox: A Generative Model for Music.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Jukebox: A Generative Model for Music

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.694978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.694978Z digest=sha256:bd60a1183e149b814e753e3b15f1b9217a3ffa07c9e9eab9ab6981f4507f9e99

Observation ef4894c0-7435-482a-88f3-22bf0f6ec3cc · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Soundstream: An end-to-end neural audio codec,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.668350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.711597Z digest=sha256:435317370a1f4eaa9a9e02f1fcc74d7439a4c846f5c88ee3401ddd3579169e0b

Observation 1a95da40-1dc8-4f35-be2e-7e394ededdab · outbound

This paper cites High Fidelity Neural Audio Compression.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music High Fidelity Neural Audio Compression

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.724873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.724873Z digest=sha256:a704b7ba2174cf44970a8125823d73d1c6d920db2110e149f9f37d1327d4fbe6

Observation d5319be5-4530-4a66-ad75-9e56b1497feb · outbound

This paper cites On generative spoken language model- ing from raw audio,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music On generative spoken language model- ing from raw audio,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.647129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.732450Z digest=sha256:7d7173f316cb4dd03e681af47d86aeb5ef4710b456f414d57d9750afaea4e277

Observation 86f30933-fecc-4be8-b275-a4d90c83ee36 · outbound

This paper cites textless-lib: a Library for Textless Spoken Language Processing.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music textless-lib: a Library for Textless Spoken Language Processing

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:00:03.069028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.738961Z digest=sha256:abf02d5655951f81cf0b7b53cc59b25be1b188eb3b20df76ba023b1c55808d93

Observation d436e966-8d7c-4ba0-a6c8-6cc1af5cefea · outbound

This paper cites MusicLM: Generating Music From Text.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music MusicLM: Generating Music From Text

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.746067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.746067Z digest=sha256:2e137e38c458514ee4da5cdda4656c6f225d1aef93ad82b1b2ea29a5d29229ec

Observation 29b3b38f-fa64-4225-bc32-ffba9023eec6 · outbound

This paper cites Simple and controllable music generation,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Simple and controllable music generation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.618553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.754180Z digest=sha256:851b716a5d5bf275b1d5d5d6091c961268f2885cefad6e0b18b22f4b26b3c3ed

Observation e5f36c6a-aafe-4c43-b936-92cdbf36fdc6 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.762282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.762282Z digest=sha256:1df9e9f3f5b8d6940a7d44d12deb8c4819524b11bf393dd054774cb0bbb5abb6

Observation 8a96e566-92fa-4a33-85b3-7a1633467b05 · outbound

This paper cites A deep learning approach for low-latency packet loss concealment of audio sig- nals in networked music performance applications,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music A deep learning approach for low-latency packet loss concealment of audio sig- nals in networked music performance applications,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.587088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.776593Z digest=sha256:b81c24cbbb355b12b50bdc25b0815e548c33afec1310b3d4158a9379b538a855

Observation f46b34ee-6b5c-4c4e-a38b-7c4fc7a83079 · outbound

This paper cites Multi-Format Contrastive Learning of Audio Representations.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Multi-Format Contrastive Learning of Audio Representations

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.786797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.786797Z digest=sha256:1e2a1f158b7fc0347c1bdc404cf4d351231290b20bd977ae4b1a04b2e7a5cc74

Observation b5095b69-910a-4388-afa3-2496334042d6 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Robust speech recognition via large-scale weak supervision,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.793373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.793373Z digest=sha256:6cfa98d0291bbc86147340933428cf940eeef377fec50667ee333551aa6d074f

Observation aaa684bb-75b4-4b82-becf-3dc6af2a1aeb · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.799420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.799420Z digest=sha256:57583895b13705686c5bf47ca1c8e9afcb02aeaed6519a5627c24ab265317d91

Observation 86d3b56c-09a3-419a-9982-f206de098694 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Distilling the Knowledge in a Neural Network

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.805950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.805950Z digest=sha256:46b81929cbb91a9810a20b24aa2315b95dba386bc5b59b876c28c333a36f1fc2

Observation abbe882d-ecfd-4aac-b51a-9a87a5514d40 · outbound

This paper cites MiniLLM: Knowledge distillation of large language models,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music MiniLLM: Knowledge distillation of large language models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.542992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.813971Z digest=sha256:253047a230a0e8f6ad653e7e6aa12a30025074b9b7fa2e3178d0c5402489c4e9

Observation 109c07dd-92fa-4297-a3bd-cefecc9a5f05 · outbound

This paper cites A simple and effective pruning approach for large lan- guage models,.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music A simple and effective pruning approach for large lan- guage models,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:00:03.497065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T15:00:02.820379Z digest=sha256:547fe646d01ef889ffd458abd1ed4b8deed0c36421df474630d73394b7023ed6

Pith citing papers

Observation b6b9e469-c7c9-4de3-8093-7581663fd5c5 · inbound

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music cites this paper.

Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T15:00:02.477825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:00:02.477825Z digest=sha256:6a6c89fab9784ddc3ab7a127108e9d7daf3f3f854437f7333e0ed1343f17ec53

Observation 72ec0687-a5c1-4e87-a590-60a3940f4dff · inbound

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models cites this paper.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.582452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.582452Z digest=sha256:16e506d8cce10d30c9c96aaac3026edce93d8a9b47d348aac868bfb22ea17fdd

Observation 0262db34-b7b9-408d-b110-e776b3ad75c3 · inbound

Breaking the Barriers of Text-Hungry and Audio-Deficient AI cites this paper.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music

Reference 130

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:29:04.385538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:28:58.216811Z digest=sha256:3a117a51519f7592a041316e80a5ec094a0b50776fc9150233ab96f3763eab6a