Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:52:01.555758Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 0 inbound Pith citation observations for arXiv:2507.09834.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:52:01.555758Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 105 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6efcc7bb-fd54-4224-a267-64c4ce3397a6 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9131fc63-07e3-43cb-bcc0-cad32f8bc346 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction CM3: A Causal Masked Multimodal Model of the Internet
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5eadcefa-5bb4-4fdf-adb0-1228793942dc · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction MusicLM: Generating Music From Text
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec3a40b2-dd06-421a-9477-24893b4722c3 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction wav2vec 2.0: A framework for self-supervised learning of speech representations
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 715f549e-44f1-4cc4-a35e-0f9fa46f1566 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Efficient self-supervised learning with contextualized target representations for vision, speech and language
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8a76069-7860-4bd9-a30a-cff8034fcfcc · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Efficient Training of Language Models to Fill in the Middle
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f9b02ea-6eba-4047-a7cf-d49fc1af0048 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction P., Whitman, B., and Lamere, P
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58732254-db4c-4a72-8c85-4cb1156ea639 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Audiolm: a language modeling approach to audio generation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a87379fc-ca28-4c30-8bb9-9a4481c3d8ea · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaf64e81-db4b-4585-b04a-93435b88f6b2 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5938745a-d9dc-4696-b80f-a42e344aa918 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Vggsound: A large-scale audio-visual dataset
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dab3047-7493-492d-8d44-2bfb8ae35f75 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Generative pretraining from pixels
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e35d0c2-c434-4d90-855a-f74abadef30e · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction W., Sutton, C., Gehrmann, S., et al
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0552ca2-b1a0-41c8-a4f0-1dcccc2dfe7f · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d30dc78-bdca-4754-a901-4879d7d66824 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction and Glass, J
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20e9f689-e8a1-4938-948c-8a3a583cc133 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Simple and controllable music generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0728646-8a08-451b-bdb5-631127642ae8 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction High fidelity neural audio compression
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d63710d3-093a-46e3-8555-74a752d6a4c4 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Audio retrieval with wavtext5k and clap training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 213abee5-84e7-458f-a2c4-21200944714c · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction and Nichol, A
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86dd9ebc-d423-403c-b9f0-432b20f2fe09 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Clotho: An audio captioning dataset
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a20690b-a4b4-4e22-a24f-7390abb4eff4 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a587f63-3347-40fb-864e-b652040bcd7f · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Taming transformers for high-resolution image synthesis
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fb8f1fe-e39f-4862-a122-58780d30709b · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Freesound datasets: a platform for the creation of open audio datasets
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10c0bb35-aae1-4d7d-8d14-3a01deb5a65f · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Fsd50k: an open dataset of human-labeled sound events
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88837d40-0691-4c63-a2e7-954918da5a11 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Incoder: A generative model for code infilling and synthesis
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efe1e0fb-4689-4be9-b974-764f913243a3 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction F., Ellis, D
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cc0151b-43c8-4f68-a72c-50d218cdacc8 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dea4d46f-e350-4a40-8e83-0d8002684e3e · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Ast: Audio spectrogram transformer
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 404fb916-edf8-46f0-8a86-0ecae1f35969 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Prompttts: Controllable text-to-speech with text descriptions
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7afb59f1-6d1d-412f-a801-0c397c5fb902 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Masked autoencoders are scalable vision learners
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fac6ee6f-7815-4a69-b397-3b8914aa50c9 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Measuring massive multitask language understanding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37e5411d-ac31-477f-b177-23cdb4519b43 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction P., Fonseca, E., Jansen, A., Liu, C., Moore, R
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f2d0011-21c9-42de-8ac8-8bd128d944be · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Denoising diffusion probabilistic models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 37702082-134f-4805-9a49-11090aa2ea46 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction A., Welbl, J., Clark, A., et al
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9b6dc69b-4214-416d-b267-fcb1c66ddf4f · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Y., Zhou, T., Wu, Y., Song, X., Song, X., and Zhou, D
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 00d3c353-1727-4733-8f28-74528ff1f58f · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a5f43fe-21fe-48cd-9b11-e7a4c61f5f59 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Masked autoencoders that listen
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 18fbb8b4-6f0f-43a0-ba38-17ce601e8167 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Multi-singer: Fast multi-singer singing voice vocoder with a large-scale corpus
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6d21ca26-1789-4528-ba53-7a17a1c28ee7 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3e716b6d-9753-4499-a314-6921d57bcdbb · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Audiogpt: Understanding and generating speech, music, sound, and talking head
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6870ffb7-d2ff-4366-8664-dc5d40ae7234 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Libri-light: A benchmark for asr with limited or no supervision
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e0b8c32a-d0ea-44be-80d5-bdb844a83846 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Scaling Laws for Neural Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16ffc977-e3ad-4f24-b9a7-43286af54204 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e9f688b1-e36f-4f42-971f-d78ebffd825b · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Fréchet audio distance: A reference-free metric for evaluating music enhancement algorithms
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1633d887-1e67-4b46-900d-bc691d61d127 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction D., Kim, B., Lee, H., and Kim, G
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e3df56f2-9a84-4b05-a932-3aa4e49424b9 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a536864-c439-427a-8f8f-8aa1e41aee2f · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction L., and Khudanpur, S
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c70c4e96-2868-49f3-a1d6-3eacbbe8deab · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 440641c6-c83a-4f55-8136-78748904b95a · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa454552-bffb-48ad-93c2-2f40b5e84d5f · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Improving text-to-audio models with synthetic captions
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a9a46066-f5de-4bcc-85f8-29c5f0fd42f4 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Audiogen: Textually guided audio generation
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation eea8ee73-f8c5-4e9f-b3ad-8eb3833668e2 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction H., Gonzalez, J., Zhang, H., and Stoica, I
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1eaaa474-480e-48ba-8ada-a5c5e5c5a144 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction BART : Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 926336f8-631a-495c-819f-45acf7f2b8d7 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Mage: Masked generative encoder to unify representation learning and image synthesis
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d4d2f52c-928c-4dd5-bf97-47024de5e9dd · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Return of unconditional generation: A self-supervised representation generation method
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7457cd7f-a098-4381-8288-37a4f9e18515 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Autoregressive Image Generation without Vector Quantization
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72d09736-9ec7-42e4-8df1-5abb5454db58 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction T., Ben-Hamu, H., Nickel, M., and Le, M
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c0b54332-2903-44e3-91be-6326942d9fcb · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction H., Le, M., Vyas, A., Shi, B., Tjandra, A., and Hsu, W.-N
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9160d2de-b07d-428b-9075-cdfdd6652322 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 58d091aa-b5d3-4e9f-856f-44b9ee5d9390 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 976a0627-d358-4905-9b8e-a3b0aeaad50c · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bc4c674-906f-4cbc-9c29-da63ae7d750b · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction WavJourney: Compositional Audio Creation with Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a19659c4-ad7d-4e37-8527-c677e4344439 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction and Hutter, F
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation abc45090-463f-488d-aa53-1b2bcc510a71 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4355582-dd0a-41b2-b185-378f2ad61988 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction and Mesaros, A
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 45847a8c-fc0d-444c-89b1-72b49a3675b1 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction D., Zou, Y., and Wang, W
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a6d23f78-8a35-4503-ad35-76443289acc7 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Autoregressive Speech Synthesis without Vector Quantization
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8492c75-26f3-4dd3-bdfc-8bf8d6648e31 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Tut database for acoustic scene classification and sound event detection
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 865789cd-22c9-45fb-b376-28e073686cea · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 427962f7-691a-4c83-bedd-0c2aed94f56c · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Representation Learning with Contrastive Predictive Coding
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f5e2073-32bc-40e7-be03-b460cf02c153 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction V oice C raft: Zero-shot speech editing and text-to-speech in the wild
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 187129eb-926d-4f26-8181-f2bebba7da45 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0db66797-635f-400b-86c0-23f335df4d04 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Efficiently scaling transformer inference
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 746df0df-790b-4155-a39d-5621ff9ff935 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Mls: A large-scale multilingual dataset for speech research
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66bfd605-536e-40f4-ae3d-798f1a1479cf · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af7feb3f-8431-4637-8538-9647bda313ee · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction High-resolution image synthesis with latent diffusion models
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed94b629-7d2c-4802-9320-835418568bf0 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Unresolved cited work
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 28dcd45f-6e60-4104-8e9c-e2b3164ada7d · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Edinburgh neural machine translation systems for wmt 16
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ef3d9cbd-4885-42fc-a0b4-6f03e0528f8e · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Aishell-3: A multi-speaker mandarin tts corpus
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfca80dc-66b3-417d-9d16-3d699de38484 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Multimodal Latent Language Modeling with Next-Token Diffusion
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2459c5e0-90cd-46b5-bdca-fca0c4b2e0b9 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Gemini: A Family of Highly Capable Multimodal Models
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9893182e-d96d-438f-ab12-5bdfbe55e876 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Givt: Generative infinite-vocabulary transformers
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b68b9619-f751-435b-bc8c-25b8e8517066 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction and Cook, P
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef59f598-6cef-4a2d-82d4-5e906935bc04 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Attention is all you need
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cca6c30e-5db6-455d-9495-cd915c13d794 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f6e0d75-48dd-4771-8145-0a7b0a03094a · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction GLUE : A multi-task benchmark and analysis platform for natural language understanding
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0e3b8bf-90c4-49c6-8219-cfe9ef22e744 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 274d4217-a588-4314-b5fe-40f706ad600b · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Not all images are worth 16x16 words: Dynamic transformers for efficient image recognition
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1455e529-5dc8-4ee6-9905-c9cc1011c562 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Opencpop: A high-quality open source chinese popular song corpus for singing voice synthesis
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73930923-0f38-4e9a-b494-22d38efcd2c9 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Audio-Agent: Leveraging LLMs For Audio Generation, Editing and Composition
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16268694-8cab-4e2e-bd06-db066041aaa6 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Diffusion models as masked autoencoders
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a2daa00a-9398-4e6a-8389-73a9985e13d2 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Next-gpt: Any-to-any multimodal llm
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 25a4ed29-91fa-4a41-8828-972a4d28a8e7 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d682d2a4-1638-47e1-b3b7-c04993354b10 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction UniAudio: An Audio Foundation Model Toward Universal Audio Generation
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bd5f09c-b538-41e3-913e-6fd813e5c2e3 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Diffsound: Discrete diffusion model for text-to-sound generation
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d8eedf5d-1025-4ae7-8e73-2b187c2b7dff · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction A Survey on Multimodal Large Language Models
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1469058c-b513-4177-8fde-bf5021854bdb · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Megabyte: Predicting million-byte sequences with multiscale transformers
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation eb91844b-c631-4bae-bcbb-c0c20c1d8031 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction and Robnik- S ikonja, M
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fb7ecedb-7c9f-4d9c-9e05-0d3081796c59 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Recurrent Neural Network Regularization
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd6314dc-4f79-4ffa-bada-bcdecbd88de3 · outbound
Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction Soundstream: An end-to-end neural audio codec
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.