Pith. sign in

Paper Citation Record · LEDGER

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction

As of 24 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 0 inbound Pith citation observations for arXiv:2507.09834.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09834 v1

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:52:01.555758Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 105 outbound references displayed

  • verified exact1
  • verified fuzzy32
  • unresolved67
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6efcc7bb-fd54-4224-a267-64c4ce3397a6 · outbound

This paper cites GPT-4 Technical Report.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:57.664745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:57.664745Z digest=sha256:001ef837a5bc1cc95416d19ebf4cac8d9f0a1533d307af3cb932e7c6fb4a27b8

Observation 9131fc63-07e3-43cb-bcc0-cad32f8bc346 · outbound

This paper cites CM3: A Causal Masked Multimodal Model of the Internet.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction CM3: A Causal Masked Multimodal Model of the Internet

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:57.735481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:57.735481Z digest=sha256:87bb6fa01789571539114f14dbc4bbf14efa60e20744862582a4aafa3f4f4a97

Observation 5eadcefa-5bb4-4fdf-adb0-1228793942dc · outbound

This paper cites MusicLM: Generating Music From Text.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction MusicLM: Generating Music From Text

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:57.869593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:57.869593Z digest=sha256:f0e7a99af9f6b99769ecd49f0923042193acc11fec34725962c4e513407789a8

Observation ec3a40b2-dd06-421a-9477-24893b4722c3 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:57.960959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:57.960959Z digest=sha256:b42fd1167474c8979f3c98d1163d45fdda793ad5400a22c463d948852ff04a1a

Observation 715f549e-44f1-4cc4-a35e-0f9fa46f1566 · outbound

This paper cites Efficient self-supervised learning with contextualized target representations for vision, speech and language.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Efficient self-supervised learning with contextualized target representations for vision, speech and language

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:58.067440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:58.067440Z digest=sha256:3c0ac0b6dd23aecfcfb166bdf789c1df250bd6d7197d9166e201ab993e917133

Observation b8a76069-7860-4bd9-a30a-cff8034fcfcc · outbound

This paper cites Efficient Training of Language Models to Fill in the Middle.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Efficient Training of Language Models to Fill in the Middle

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:58.202849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:58.202849Z digest=sha256:9e28dce373c72aadcb52977e002def7f3572c8550dea7ae6ef70e30c917899ae

Observation 7f9b02ea-6eba-4047-a7cf-d49fc1af0048 · outbound

This paper cites P., Whitman, B., and Lamere, P.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction P., Whitman, B., and Lamere, P

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:58.295858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:58.295858Z digest=sha256:f4a30f331bc4eb34f082103de37f8cdbc8b9e03d1afa6bea0af401aa33f5d9a8

Observation 58732254-db4c-4a72-8c85-4cb1156ea639 · outbound

This paper cites Audiolm: a language modeling approach to audio generation.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Audiolm: a language modeling approach to audio generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:58.418812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:58.418812Z digest=sha256:05e784515e3c023095ecdee3b7a65ad3edb6d695fc792578021a3868d3b3f5f8

Observation a87379fc-ca28-4c30-8bb9-9a4481c3d8ea · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:58.544942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:58.544942Z digest=sha256:ec00d1e9ad62d9871d891a7c52302c5d269f8813786a26d7a857e7fd8af2eaec

Observation aaf64e81-db4b-4585-b04a-93435b88f6b2 · outbound

This paper cites an unresolved cited work.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:58.639495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:58.639495Z digest=sha256:da914b965e78cf0a286290787c2036ac16e21c0176825e80278b658cd7c6314b

Observation 5938745a-d9dc-4696-b80f-a42e344aa918 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Vggsound: A large-scale audio-visual dataset

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:58.766911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:58.766911Z digest=sha256:51549f669a825b1a72e206276b2608a358923381a7be0c2270f4083ac5cf1f52

Observation 0dab3047-7493-492d-8d44-2bfb8ae35f75 · outbound

This paper cites Generative pretraining from pixels.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Generative pretraining from pixels

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:58.854783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:58.854783Z digest=sha256:e3e21d98c77136757f09fdda3b5f76fafb28d57bda92bd34c5fd455e1d81c7e7

Observation 7e35d0c2-c434-4d90-855a-f74abadef30e · outbound

This paper cites W., Sutton, C., Gehrmann, S., et al.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction W., Sutton, C., Gehrmann, S., et al

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:58.920337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:58.920337Z digest=sha256:ff1e0b8c7edcf823b4263a30f2cc1cf41a62cd26785ccc21d1ecefe82b857c28

Observation e0552ca2-b1a0-41c8-a4f0-1dcccc2dfe7f · outbound

This paper cites W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:58.987922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:58.987922Z digest=sha256:b30228bbee97a240b7ec31204359f05994d055d2848cb0cfc89dcad1bb39aed9

Observation 7d30dc78-bdca-4754-a901-4879d7d66824 · outbound

This paper cites and Glass, J.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction and Glass, J

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:59.056092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:59.056092Z digest=sha256:e6798174c6af156f51835d5520e4ef46d30e386c9a4ebd30e9c5c18ac5018d5b

Observation 20e9f689-e8a1-4938-948c-8a3a583cc133 · outbound

This paper cites Simple and controllable music generation.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Simple and controllable music generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:59.118450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:59.118450Z digest=sha256:bf3e5bd1147e8fbb48523cf9388b656b63a8ba13d7085097f4817198c6490785

Observation c0728646-8a08-451b-bdb5-631127642ae8 · outbound

This paper cites High fidelity neural audio compression.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction High fidelity neural audio compression

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:59.214401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:59.214401Z digest=sha256:51660f796213c8ca29bf04f1d553143da0923b421d91b2545ea61cfe01987324

Observation d63710d3-093a-46e3-8555-74a752d6a4c4 · outbound

This paper cites Audio retrieval with wavtext5k and clap training.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Audio retrieval with wavtext5k and clap training

Reference 18

Resolution
verified exact
doi, observed 2026-08-06T17:52:02.088576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:51:59.310297Z digest=sha256:695de4a9602557fee0394e79489388d50eae748950e1969ba2268f3801370e34

Observation 213abee5-84e7-458f-a2c4-21200944714c · outbound

This paper cites and Nichol, A.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction and Nichol, A

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:59.385684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:59.385684Z digest=sha256:d14818a138def005190f48ca4e3bb5085ee286d2155ff24d28789af6433431c0

Observation 86dd9ebc-d423-403c-b9f0-432b20f2fe09 · outbound

This paper cites Clotho: An audio captioning dataset.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Clotho: An audio captioning dataset

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:59.490738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:59.490738Z digest=sha256:46b58170db5dc29dc1acc2d95177e44ad5a7464da6772bbf9f26ace6f715ec6b

Observation 5a20690b-a4b4-4e22-a24f-7390abb4eff4 · outbound

This paper cites M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:59.563010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:59.563010Z digest=sha256:ee720fe73679298e4d6642229c8dbf83e9640b3adbb1538d162039a25db21412

Observation 6a587f63-3347-40fb-864e-b652040bcd7f · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Taming transformers for high-resolution image synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:59.648011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:59.648011Z digest=sha256:d8139964f1a20808e0c951cb774c5f093122d3be0cbee14086b20b38cd0ac088

Observation 5fb8f1fe-e39f-4862-a122-58780d30709b · outbound

This paper cites Freesound datasets: a platform for the creation of open audio datasets.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Freesound datasets: a platform for the creation of open audio datasets

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:59.785341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:59.785341Z digest=sha256:b71a25dd5cd654dd54aaf889694de3a88fa44bc07acb160c5e78074652eccffe

Observation 10c0bb35-aae1-4d7d-8d14-3a01deb5a65f · outbound

This paper cites Fsd50k: an open dataset of human-labeled sound events.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Fsd50k: an open dataset of human-labeled sound events

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:59.829775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:59.829775Z digest=sha256:976162f2a320ea7b39c6ec8ce55443ab1dc61267e0d04a86d53cecad47fd35a3

Observation 88837d40-0691-4c63-a2e7-954918da5a11 · outbound

This paper cites Incoder: A generative model for code infilling and synthesis.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Incoder: A generative model for code infilling and synthesis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:59.937178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:59.937178Z digest=sha256:bee598eb5b6576d8070a1e16a5ae89140bd3397f21506ea16d87017868380c7a

Observation efe1e0fb-4689-4be9-b974-764f913243a3 · outbound

This paper cites F., Ellis, D.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction F., Ellis, D

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:00.050091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:00.050091Z digest=sha256:0f043d11f5bdde1cc9ded8f2aaa4c49e89112187fcc1bc341596cd9f35c9675f

Observation 8cc0151b-43c8-4f68-a72c-50d218cdacc8 · outbound

This paper cites Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:00.192473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:00.192473Z digest=sha256:072edcab1604f9e5b21c8d2548eca7aba3fc84c5e3321f4e3979bcb2cb31a017

Observation dea4d46f-e350-4a40-8e83-0d8002684e3e · outbound

This paper cites Ast: Audio spectrogram transformer.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Ast: Audio spectrogram transformer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:00.328565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:00.328565Z digest=sha256:abe31cc09e5b4b47f08ea0802b0a4b7b951a72194362bf1466b565e915f10e01

Observation 404fb916-edf8-46f0-8a86-0ecae1f35969 · outbound

This paper cites Prompttts: Controllable text-to-speech with text descriptions.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Prompttts: Controllable text-to-speech with text descriptions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:00.412842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:00.412842Z digest=sha256:b57a709826c04a967a1f42820b48aa8f18dcfad265114c9804e44391956df1eb

Observation 7afb59f1-6d1d-412f-a801-0c397c5fb902 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Masked autoencoders are scalable vision learners

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:00.547835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:00.547835Z digest=sha256:3f3e7fe7a42ab6c88a887a201c775db51305ca98ccdc1c077c0ad329337bddcf

Observation fac6ee6f-7815-4a69-b397-3b8914aa50c9 · outbound

This paper cites Measuring massive multitask language understanding.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Measuring massive multitask language understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:00.682210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:00.682210Z digest=sha256:d62b5aa8d7baf74461f70976c26443b05f9b788862701944abe8c297aa6eccab

Observation 37e5411d-ac31-477f-b177-23cdb4519b43 · outbound

This paper cites P., Fonseca, E., Jansen, A., Liu, C., Moore, R.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction P., Fonseca, E., Jansen, A., Liu, C., Moore, R

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:00.809190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:00.809190Z digest=sha256:c6f90627c026499bb11e23154d44c8df268b2e30fb870b15bcc9428e6e9df681

Observation 2f2d0011-21c9-42de-8ac8-8bd128d944be · outbound

This paper cites Denoising diffusion probabilistic models.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Denoising diffusion probabilistic models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.659988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:00.955163Z digest=sha256:f8d9925ba0e6d23410a655b71c32ea3fbd703502b6320c077020de55e89930fb

Observation 37702082-134f-4805-9a49-11090aa2ea46 · outbound

This paper cites A., Welbl, J., Clark, A., et al.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction A., Welbl, J., Clark, A., et al

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.633753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.045071Z digest=sha256:071f766ca29561033cb5deab5c746c246de1aca5283395b470911a59311d0304

Observation 9b6dc69b-4214-416d-b267-fcb1c66ddf4f · outbound

This paper cites Y., Zhou, T., Wu, Y., Song, X., Song, X., and Zhou, D.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Y., Zhou, T., Wu, Y., Song, X., Song, X., and Zhou, D

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.608301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.174452Z digest=sha256:b0b3404e682c5595fe3c48c92ad39b6f4065546dfe74f7db013efe9c2e93030a

Observation 00d3c353-1727-4733-8f28-74528ff1f58f · outbound

This paper cites Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.184215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.184215Z digest=sha256:90549c069ec93d472ccc78d76d9c4a55075e62fa126f1beafa5be5ac8873f682

Observation 2a5f43fe-21fe-48cd-9b11-e7a4c61f5f59 · outbound

This paper cites Masked autoencoders that listen.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Masked autoencoders that listen

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.587728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.213931Z digest=sha256:78f5e76e679532b14c55c1f4232d8eebab06ff064e55dad9c9307b9518a73609

Observation 18fbb8b4-6f0f-43a0-ba38-17ce601e8167 · outbound

This paper cites Multi-singer: Fast multi-singer singing voice vocoder with a large-scale corpus.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Multi-singer: Fast multi-singer singing voice vocoder with a large-scale corpus

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.564211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.219238Z digest=sha256:26824afc69a579d4f402dd9906943e330a78b4722ed33d656e76ca6a4e78edc9

Observation 6d21ca26-1789-4528-ba53-7a17a1c28ee7 · outbound

This paper cites Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.539545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.223616Z digest=sha256:ee75f70559a0c1c96416d3125a296ab9a10e9acea9e22e1cb763ce272d684ffe

Observation 3e716b6d-9753-4499-a314-6921d57bcdbb · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Audiogpt: Understanding and generating speech, music, sound, and talking head

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.517952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.228181Z digest=sha256:b31ff2a179d4bbafca6e5c3655a27bac00ec64539444e1a11a555dd5557d1a2c

Observation 6870ffb7-d2ff-4366-8664-dc5d40ae7234 · outbound

This paper cites Libri-light: A benchmark for asr with limited or no supervision.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Libri-light: A benchmark for asr with limited or no supervision

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.498114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.232203Z digest=sha256:9f45d67bb60415f2d7db3ebe9e8a2f86aa34793249863abdeafd1d6f5024bcd6

Observation e0b8c32a-d0ea-44be-80d5-bdb844a83846 · outbound

This paper cites Scaling Laws for Neural Language Models.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Scaling Laws for Neural Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.237816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.237816Z digest=sha256:31cf75579293feebfd6770b6d1bf36de52064c1f4b843d08d35ce8ee23722630

Observation 16ffc977-e3ad-4f24-b9a7-43286af54204 · outbound

This paper cites an unresolved cited work.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:52:03.480310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.243761Z digest=sha256:01558bca483d1be3fc9f2fe9af710cfd26156d6cec69da040de570d5ce2cc7fa

Observation e9f688b1-e36f-4f42-971f-d78ebffd825b · outbound

This paper cites Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.248336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.248336Z digest=sha256:d08d3ac6ec7c2af8ee2f61bdfd325f6550e7a73bfa521251f824cd1fe45edd4b

Observation 1633d887-1e67-4b46-900d-bc691d61d127 · outbound

This paper cites D., Kim, B., Lee, H., and Kim, G.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction D., Kim, B., Lee, H., and Kim, G

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.456885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.253604Z digest=sha256:1b36e816b46fc2aa27a9bd478be0bc86b653abe7c46a5636a37695f74e9158c4

Observation e3df56f2-9a84-4b05-a932-3aa4e49424b9 · outbound

This paper cites an unresolved cited work.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.258735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.258735Z digest=sha256:d0112c8caa24f178d432fd879bedff79667388a6359bd1df039efc19f6a06db9

Observation 8a536864-c439-427a-8f8f-8aa1e41aee2f · outbound

This paper cites L., and Khudanpur, S.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction L., and Khudanpur, S

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.421191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.263424Z digest=sha256:2d18b9f530f8432fe7f7591ff3877b2bc2da6da73fc6aa327d74177f89f11f49

Observation c70c4e96-2868-49f3-a1d6-3eacbbe8deab · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.402332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.267953Z digest=sha256:82ad63c425ce308e32fb381596a47a79269f1bbfd5f5afdf8ea4e95585ba4f8b

Observation 440641c6-c83a-4f55-8136-78748904b95a · outbound

This paper cites Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.273007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.273007Z digest=sha256:e25572f5898d96b3ab10e7a0d027b25f94d4e000ed17d337326091b0e18b2340

Observation fa454552-bffb-48ad-93c2-2f40b5e84d5f · outbound

This paper cites Improving text-to-audio models with synthetic captions.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Improving text-to-audio models with synthetic captions

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.383749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.279126Z digest=sha256:e7d5348065d9ddf5e7cc9b8f1fe5138d81b95e4b4a2b9b75d1cca8304278d62c

Observation a9a46066-f5de-4bcc-85f8-29c5f0fd42f4 · outbound

This paper cites Audiogen: Textually guided audio generation.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Audiogen: Textually guided audio generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.360536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.284072Z digest=sha256:3524b4d2f1feba400eb6045e426043aa58176707561e855dc909b8de5db9e280

Observation eea8ee73-f8c5-4e9f-b3ad-8eb3833668e2 · outbound

This paper cites H., Gonzalez, J., Zhang, H., and Stoica, I.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction H., Gonzalez, J., Zhang, H., and Stoica, I

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.341394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.288125Z digest=sha256:3f278e38454859e7901f155871445e1e33e69a77d31d9386f293618376628bde

Observation 1eaaa474-480e-48ba-8ada-a5c5e5c5a144 · outbound

This paper cites BART : Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction BART : Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.292238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.292238Z digest=sha256:8350909e29fb89a19c7768e5a2cd3a59108ebb8282558798ec2de3c11b105da3

Observation 926336f8-631a-495c-819f-45acf7f2b8d7 · outbound

This paper cites Mage: Masked generative encoder to unify representation learning and image synthesis.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Mage: Masked generative encoder to unify representation learning and image synthesis

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.321304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.296437Z digest=sha256:2cd1eb369250c1f85c712cf3d55a94e4607026c9e7262b8425a4400b5fd35918

Observation d4d2f52c-928c-4dd5-bf97-47024de5e9dd · outbound

This paper cites Return of unconditional generation: A self-supervised representation generation method.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Return of unconditional generation: A self-supervised representation generation method

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.305695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.301026Z digest=sha256:b590a0da789b5cf634c43d43d129f1a5db0edee2c33df88ecc0752c80a6eb617

Observation 7457cd7f-a098-4381-8288-37a4f9e18515 · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Autoregressive Image Generation without Vector Quantization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.305060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.305060Z digest=sha256:a84f142696c47a206092814879344758769197e02d5cc5800fadd6d0283f047a

Observation 72d09736-9ec7-42e4-8df1-5abb5454db58 · outbound

This paper cites T., Ben-Hamu, H., Nickel, M., and Le, M.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction T., Ben-Hamu, H., Nickel, M., and Le, M

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.287718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.309511Z digest=sha256:cadfec3a7ff72d7dd88dec3702c0159173c562d52a0cd400db61a0a415f670dd

Observation c0b54332-2903-44e3-91be-6326942d9fcb · outbound

This paper cites H., Le, M., Vyas, A., Shi, B., Tjandra, A., and Hsu, W.-N.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction H., Le, M., Vyas, A., Shi, B., Tjandra, A., and Hsu, W.-N

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.271019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.313858Z digest=sha256:fd5f1bd205b99f6ebe1ab74bffb403672646b15f2552177e2fe8ce01f205c2de

Observation 9160d2de-b07d-428b-9075-cdfdd6652322 · outbound

This paper cites an unresolved cited work.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:52:03.254397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.319281Z digest=sha256:112824b438ff39c2078a9afd5553ab030c9b2fb2e2c321e04e1d8642814c128e

Observation 58d091aa-b5d3-4e9f-856f-44b9ee5d9390 · outbound

This paper cites an unresolved cited work.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:52:03.235176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.323474Z digest=sha256:8029577851d0f28952eb95d2594ffbe41db9161426183b37856b4be8d77cdfef

Observation 976a0627-d358-4905-9b8e-a3b0aeaad50c · outbound

This paper cites an unresolved cited work.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.328026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.328026Z digest=sha256:3748a7ee555e832c36219e67e49f777beb03341ce6dd0505dbc87e3349483cb0

Observation 6bc4c674-906f-4cbc-9c29-da63ae7d750b · outbound

This paper cites WavJourney: Compositional Audio Creation with Large Language Models.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction WavJourney: Compositional Audio Creation with Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.332297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.332297Z digest=sha256:afa942a340e7e6402f6467ffee6f1c0ef1c449d822c7f5f728d2f6a220963f40

Observation a19659c4-ad7d-4e37-8527-c677e4344439 · outbound

This paper cites and Hutter, F.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction and Hutter, F

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.218694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.336968Z digest=sha256:555c7afc1fe0e2710e864fff896b68c8446d57d0cd3d851ce59653ee9fdcb154

Observation abc45090-463f-488d-aa53-1b2bcc510a71 · outbound

This paper cites Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.341231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.341231Z digest=sha256:3616450148f5087b1071f9b880149381a3bc4653788d0fdd0cbf12e1007b803c

Observation f4355582-dd0a-41b2-b185-378f2ad61988 · outbound

This paper cites and Mesaros, A.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction and Mesaros, A

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.198657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.345701Z digest=sha256:396208b1a3bb24c2e7a38fd217d34caec587dad898c09ce4b7ad48303a1da91b

Observation 45847a8c-fc0d-444c-89b1-72b49a3675b1 · outbound

This paper cites D., Zou, Y., and Wang, W.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction D., Zou, Y., and Wang, W

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.159641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.350124Z digest=sha256:5d92cb6540086665e1afd5d7e330b89f3a49938abbda0516d81d7b64a9824dc5

Observation a6d23f78-8a35-4503-ad35-76443289acc7 · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Autoregressive Speech Synthesis without Vector Quantization

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.354517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.354517Z digest=sha256:a07945655c48d5244d8dd2bf6a9e06e2d779209e6903760787f7b6257516f835

Observation f8492c75-26f3-4dd3-bdfc-8bf8d6648e31 · outbound

This paper cites Tut database for acoustic scene classification and sound event detection.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Tut database for acoustic scene classification and sound event detection

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.136505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.358779Z digest=sha256:c320a48ccbbf5fa411e99fdc7f2895aaacfa90a68148ed86ea5c42d2b3803b4e

Observation 865789cd-22c9-45fb-b376-28e073686cea · outbound

This paper cites an unresolved cited work.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:52:03.115867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.364310Z digest=sha256:bdfb267f58c6a2cb4d5ebd649fdece7e5153a5dbd39597285f734b6513b05171

Observation 427962f7-691a-4c83-bedd-0c2aed94f56c · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Representation Learning with Contrastive Predictive Coding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.368588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.368588Z digest=sha256:c3d0af8dedf881bbb4dc2a7b8f615804a62c651487fb2c886f0374f71bc01abd

Observation 3f5e2073-32bc-40e7-be03-b460cf02c153 · outbound

This paper cites V oice C raft: Zero-shot speech editing and text-to-speech in the wild.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction V oice C raft: Zero-shot speech editing and text-to-speech in the wild

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.373095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.373095Z digest=sha256:c85ea7ff049b3719152bb85bdf70fd584ae842460ce0f625b4e5f2b90c834612

Observation 187129eb-926d-4f26-8181-f2bebba7da45 · outbound

This paper cites an unresolved cited work.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:52:03.095751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.377763Z digest=sha256:b4aea6c26a999f45b1ec9a6b55f84c20142422eb565e7c84d1e581d69edb8778

Observation 0db66797-635f-400b-86c0-23f335df4d04 · outbound

This paper cites Efficiently scaling transformer inference.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Efficiently scaling transformer inference

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:03.077930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.382718Z digest=sha256:d7698419e3b280f876328fa7735f7c24a0c57f7563f9e45cf53789bfc2761a59

Observation 746df0df-790b-4155-a39d-5621ff9ff935 · outbound

This paper cites Mls: A large-scale multilingual dataset for speech research.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Mls: A large-scale multilingual dataset for speech research

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.387418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.387418Z digest=sha256:bd528a669c264c9e021b62fde39d8d4be9516e4783726b444410287de42885df

Observation 66bfd605-536e-40f4-ae3d-798f1a1479cf · outbound

This paper cites an unresolved cited work.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.392376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.392376Z digest=sha256:0622610b67d4da757a7774b6ce74ca98f96a5b0ca7b617888bc5f70a069afd7c

Observation af7feb3f-8431-4637-8538-9647bda313ee · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction High-resolution image synthesis with latent diffusion models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.396597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.396597Z digest=sha256:6d9ce55fe6d1c2afb77e7c6885561d97f815514d83dcbc72c58e5cc5a22f4359

Observation ed94b629-7d2c-4802-9320-835418568bf0 · outbound

This paper cites an unresolved cited work.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:52:03.009050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.401383Z digest=sha256:a75ce504a98af1a7e4693b843f1260637cf77a05c3ce3b58f5c0efbdad3e6847

Observation 28dcd45f-6e60-4104-8e9c-e2b3164ada7d · outbound

This paper cites Edinburgh neural machine translation systems for wmt 16.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Edinburgh neural machine translation systems for wmt 16

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:02.980768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.408251Z digest=sha256:0c77c0ef1f3590a9454c49cf7f13a5feb9d7d4fc4b576ce28e36b4efecefa4ad

Observation ef3d9cbd-4885-42fc-a0b4-6f03e0528f8e · outbound

This paper cites Aishell-3: A multi-speaker mandarin tts corpus.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Aishell-3: A multi-speaker mandarin tts corpus

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.413702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.413702Z digest=sha256:2d225f8d8c0b45bbbe068ff0ee9d890bbede064a43b03ec447b1aed08fb25d9c

Observation bfca80dc-66b3-417d-9d16-3d699de38484 · outbound

This paper cites Multimodal Latent Language Modeling with Next-Token Diffusion.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Multimodal Latent Language Modeling with Next-Token Diffusion

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.425532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.425532Z digest=sha256:cac3fa896f9cf219fcdd277bf4a6d8dafd23ec22a4c810b529b825b86e1f9a57

Observation 2459c5e0-90cd-46b5-bdca-fca0c4b2e0b9 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Gemini: A Family of Highly Capable Multimodal Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.430636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.430636Z digest=sha256:681d06dee5a1bc7d753148ed58fc914e8407bc703b1936170c499cdc2447625a

Observation 9893182e-d96d-438f-ab12-5bdfbe55e876 · outbound

This paper cites Givt: Generative infinite-vocabulary transformers.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Givt: Generative infinite-vocabulary transformers

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:02.950079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.436170Z digest=sha256:4324b32f3277aa580eb9acdf9dc5b63e56ac63d100fa511446e1a2650e333989

Observation b68b9619-f751-435b-bc8c-25b8e8517066 · outbound

This paper cites and Cook, P.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction and Cook, P

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.440950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.440950Z digest=sha256:9c042dd80bfc365527ea6d2f4e6e4b0206ff194c8c5edb0939f1fb39fa22b0c2

Observation ef59f598-6cef-4a2d-82d4-5e906935bc04 · outbound

This paper cites Attention is all you need.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Attention is all you need

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.445811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.445811Z digest=sha256:7aec2b986d560a281c0e5595845658635ee70ddf5442944f7c84a49eb39ab0c6

Observation cca6c30e-5db6-455d-9495-cd915c13d794 · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.451574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.451574Z digest=sha256:e82ad96cd6b506eabbbbee1d1cb4738272490c4116a0ef5418ffc716358af059

Observation 5f6e0d75-48dd-4771-8145-0a7b0a03094a · outbound

This paper cites GLUE : A multi-task benchmark and analysis platform for natural language understanding.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction GLUE : A multi-task benchmark and analysis platform for natural language understanding

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.457448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.457448Z digest=sha256:2c1767b2de3d2110f13bc43921e202be6d515593bb394aed812e0eb67a117574

Observation d0e3b8bf-90c4-49c6-8219-cfe9ef22e744 · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.469864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.469864Z digest=sha256:a126a7339d7efbb05d56ada9844cd03b37a2a7ee54f414d99e2929aae2038d1e

Observation 274d4217-a588-4314-b5fe-40f706ad600b · outbound

This paper cites Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:02.882459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.479739Z digest=sha256:8e4779849dd9bbec02789e17c63dc3ba76cad66eb275161b52d8145e1697722a

Observation 1455e529-5dc8-4ee6-9905-c9cc1011c562 · outbound

This paper cites Opencpop: A high-quality open source chinese popular song corpus for singing voice synthesis.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Opencpop: A high-quality open source chinese popular song corpus for singing voice synthesis

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.485978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.485978Z digest=sha256:1ea8bd123794bfe80cbce87d578f509183c427f64afeb3d806d6eda11be1635b

Observation 73930923-0f38-4e9a-b494-22d38efcd2c9 · outbound

This paper cites Audio-Agent: Leveraging LLMs For Audio Generation, Editing and Composition.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Audio-Agent: Leveraging LLMs For Audio Generation, Editing and Composition

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.491237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.491237Z digest=sha256:281be56d5eafa6d03508df3b3ef0c0c9e1cbe08386d32870e4f63fa3e6030434

Observation 16268694-8cab-4e2e-bd06-db066041aaa6 · outbound

This paper cites Diffusion models as masked autoencoders.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Diffusion models as masked autoencoders

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:02.856725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.497486Z digest=sha256:aa26f213d0f9a3ccd3b62102a0118da4c5de65553e484a9c0414a238235d824e

Observation a2daa00a-9398-4e6a-8389-73a9985e13d2 · outbound

This paper cites Next-gpt: Any-to-any multimodal llm.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Next-gpt: Any-to-any multimodal llm

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:02.832100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.508042Z digest=sha256:fcfbc220ab362dd2eb9c45dd215cea806c8f51c2e458a797794e8c336ae21d94

Observation 25a4ed29-91fa-4a41-8828-972a4d28a8e7 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:02.811159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.515390Z digest=sha256:c6a60db40cba3d2b76ce2f5dfe9ce4ad888374c06f3906e722c837df924105e8

Observation d682d2a4-1638-47e1-b3b7-c04993354b10 · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.524187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.524187Z digest=sha256:bf6db99636ece0e97b7a33da82e3c56d06e18be60f7dba012a5cb35c592c1fcb

Observation 2bd5f09c-b538-41e3-913e-6fd813e5c2e3 · outbound

This paper cites Diffsound: Discrete diffusion model for text-to-sound generation.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Diffsound: Discrete diffusion model for text-to-sound generation

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:02.787593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.531566Z digest=sha256:c53f3e024e706dbecb356b0f5526594a61d47ba194440545204247850383f1f9

Observation d8eedf5d-1025-4ae7-8e73-2b187c2b7dff · outbound

This paper cites A Survey on Multimodal Large Language Models.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction A Survey on Multimodal Large Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.536164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.536164Z digest=sha256:ba81ea696b0678730cbeadd14388c03eab42f793dbfc396f53c0c694da013930

Observation 1469058c-b513-4177-8fde-bf5021854bdb · outbound

This paper cites Megabyte: Predicting million-byte sequences with multiscale transformers.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Megabyte: Predicting million-byte sequences with multiscale transformers

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:02.766346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.540872Z digest=sha256:af2b30a6d23f19d519ebe588f252b4f893a0d53daba66763799f8194d93bc8ed

Observation eb91844b-c631-4bae-bcbb-c0c20c1d8031 · outbound

This paper cites and Robnik- S ikonja, M.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction and Robnik- S ikonja, M

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:52:02.743229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-06T17:52:01.545972Z digest=sha256:8ed982c20377a294edad5ab41758f53a7d8a71bdb20ed9ed823a6fba145328da

Observation fb7ecedb-7c9f-4d9c-9e05-0d3081796c59 · outbound

This paper cites Recurrent Neural Network Regularization.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Recurrent Neural Network Regularization

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.550702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.550702Z digest=sha256:56ac576260989ad79871f14e4ee2d9735e84854c87d3a9301377debb58ead745

Observation bd6314dc-4f79-4ffa-bada-bcdecbd88de3 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Soundstream: An end-to-end neural audio codec

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:01.555758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:01.555758Z digest=sha256:662d3b1d768c694713b6094493e3bfad57c3e2974aad36409e93d659cfede113

Pith citing papers

No inbound Pith citation observations are available.