Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:12:24.123186Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 4 inbound Pith citation observations for arXiv:2412.04917.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:12:24.123186Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:03:30.685319Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T20:45:08.132130Z
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 57e20474-ba2e-4d9a-90ac-ef77d669ecd1 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b603eacb-3b43-4ded-94de-e74e081eb72b · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Ichigo: Mixed-Modal Early-Fusion Realtime Voice Assistant
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8127f057-5242-4aba-a1d2-86ff74d7fb31 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Moshi: a speech-text foundation model for real-time dialogue
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3021b48b-e7f1-4a94-9036-9dd1c853f500 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c56b5a3-535f-4f59-9386-a4c401157592 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5267b4e2-7a21-48c7-9c16-7b84feba1748 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 946cfcaf-f1c7-4e7f-af8e-a168e8f7037a · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners VITA: Towards Open-Source Interactive Omni Multimodal LLM
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e74ffe0d-4a08-45b2-8b09-ca0d96878287 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b50dc08f-184c-43ab-8d88-a206a844f672 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners GPT-4o System Card
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a03dbb64-8ad8-47ff-8b52-f7946be435e7 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners WavChat: A Survey of Spoken Dialogue Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7829fab8-35e1-49b8-b2aa-881cca0d32f5 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 979a1fc0-da93-48da-afcc-623cc445ad85 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Making Flow-Matching-Based Zero-Shot Text-to-Speech Laugh as You Like
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1f51dc1-1cfa-4c52-870c-2446676844fe · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1aff6de-b716-4a8b-b614-80ac8f879a81 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners MoBoAligner: a Neural Alignment Model for Non-autoregressive TTS with Monotonic Boundary Search
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7867242-133a-4905-b1b3-63cb5db9de4e · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74d342f8-f9e7-43bc-902d-c2e766602add · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners DelightfulTTS 2: End-to-End Speech Synthesis with Adversarial Vector-Quantized Auto-Encoders
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3caa5dde-4d75-4246-a896-2f262d2f30f0 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Autoregressive Speech Synthesis without Vector Quantization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d201cb1d-f8ce-4341-9bf9-a7d36e2a74ea · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Spoken Question Answering and Speech Continuation Using Spectrogram-Powered LLM
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69168993-6dbb-403d-8c7b-dc73dccc4d96 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Mo\^usai: Text-to-Music Generation with Long-Context Latent Diffusion
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f07153a0-ea82-4e04-9ff0-c196a1530593 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Bespoke Solvers for Generative Flow Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a15dfbd5-666e-4a52-a256-a949d4b516ca · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d06573db-0160-4c6a-98c2-0bab659855e6 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Improving and generalizing flow-based generative models with minibatch optimal transport
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e440572-9239-4a3b-b3b3-366bc5f11d2b · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72795f17-3505-4ec4-b58a-f73569adbbd4 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44456515-7b16-4442-ad1d-e431217f467f · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ced33df-bf4e-4c94-92db-cee7a6a5191c · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33d97c09-3a17-4b48-a6a0-28313112caba · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Qwen2 Technical Report
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ad4dfac-b493-4042-948b-b91b33bbdc11 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24fb27b1-dca9-4f65-8ef2-34f80ba7ec56 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Boosting Diffusion Model for Spectrogram Up-sampling in Text-to-speech: An Empirical Study
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57606263-1fd7-4937-bb54-c3495b0086f1 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21dc1169-1c22-44d0-b05a-abea892a15de · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners PromptTTS 2: Describing and Generating Voices with Text Prompt
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11b02b08-2e59-4e24-bd95-2b7283f4dcf8 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bab0be1d-6279-4510-8f4c-9b040b3495a8 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners MarDini: Masked Autoregressive Diffusion for Video Generation at Scale
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ec7112d-9787-477e-8730-7f7185ce16b4 · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Understanding DDPM Latent Codes Through Optimal Transport
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cce992c9-8d14-4e91-b54e-8938282f5b6a · outbound
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners Recent Advances in Speech Language Models: A Survey
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f425dfd-1a03-4aa3-8f97-c1519f08c345 · inbound
On The Landscape of Spoken Language Models: A Comprehensive Survey Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a83a50da-cb83-4dd3-a6cc-78708deb9984 · inbound
IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57703305-fd3c-4bd1-8837-dd90a5334c5e · inbound
Next Tokens Denoising for Speech Synthesis Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebd68a4c-3872-4398-9419-19b1b8df4f0f · inbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.