Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:54:19.340049Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 2 inbound Pith citation observations for arXiv:2507.20091.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:54:19.340049Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-11T00:49:26.507281Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T19:46:15.325325Z
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dee9aa13-00b1-4644-b147-bbf0eba5df4f · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d4af0ee-20f8-4091-a709-a63d1d3f8468 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Dm-codec: Distilling multimodal representations for speech tokenization
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c92597f9-8e37-453a-b3f5-2f29d8379e52 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models dMel: Speech Tokenization made Simple
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dddf75bb-7c71-40a6-9544-69c2e8c066c2 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Audiolm: a language modeling approach to audio generation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d818c27-e966-48ff-85f2-dccce0497ae0 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models SoundStorm: Efficient Parallel Audio Generation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e085d97-ffc2-4cef-b51c-8f5fe7d5dae0 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Giveness, contrasitiveness, definiteness, subjects, topics, and point of view
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3d1e0e8c-5558-49b9-85ea-5876e630df6e · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccb609e8-ddf5-4482-843d-273d1b8e377a · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models DrawSpeech: Expressive Speech Synthesis Using Prosodic Sketches as Control Conditions
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24823737-1366-4758-af51-1ecf7adb7e91 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 713a6d3e-8e89-452f-9f2a-83960024e07b · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models High Fidelity Neural Audio Compression
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fccdd45-b965-454c-8f5b-97c77fc1e9f2 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Moshi: a speech-text foundation model for real-time dialogue
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e274388a-c2c9-48b9-8754-7b5b1eedbe3c · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Elevenlabs voice generation platform
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5612812a-b60a-4f0a-9d37-dae60a462b24 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Recent advances in discrete speech tokens: A review
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0d96cdc-cac7-4247-9e64-4290d0ad8f7d · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Lora: Low-rank adaptation of large language models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 285c13b6-7cfc-4155-a2fc-72d0f0da6136 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2f588c9-eb8b-47b3-8819-c41769fd9cdc · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models RepCodec: A Speech Representation Codec for Speech Tokenization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73502204-a433-4ace-8bee-40ce7176eac5 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Crossing the uncanny valley of conversational voice, 2025
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 357c6328-21d4-44b9-a094-d50b74a4d917 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models An open source emotional speech corpus for human robot interaction applications
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e25b62d3-fec4-4b4c-aa00-4d0ed08db09c · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Style Mixture of Experts for Expressive Text-To-Speech Synthesis
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c2ec780-f37a-4266-a023-2e9471ff796d · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7da3462-7686-40ce-81f2-9586298d6314 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Libri-light: A benchmark for asr with limited or no supervision
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8995600a-5152-4815-b6a7-4fd51b794054 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d243f73c-5256-410b-8999-06111fd74af7 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models On generative spoken language modeling from raw audio
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6e6aa01-e127-43e4-b570-f2ed7cc28c2f · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Whisma: A speech-llm to perform zero-shot spoken language understanding
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f5543cb5-b5ab-4813-95be-6a0a387dfe0b · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dae726de-a84f-464e-a818-55039915a068 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Generative spoken dialogue language modeling
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f05a6762-2cf0-4e7b-a310-5d3613d603ee · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Spirit-lm: Interleaved spoken and written language model
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1da4e900-f7a8-4349-9c57-8b093de1d1f8 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Long-Form Speech Generation with Spoken Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9db9bdc-d2cf-4d6a-bbd2-6002769aedf8 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Robust speech recognition via large-scale weak supervision
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 912c46e5-47a8-4262-a262-cbcd8c6d09ab · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Prosospeech: Enhancing prosody with quantized vector pre-training in text-to-speech
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8c198c27-1d04-46e8-90d3-e93b341966b3 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models AudioPaLM: A Large Language Model That Can Speak and Listen
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57f1eb8c-049b-4b88-b50c-66b3bf03ffe1 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Shechtman, S
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6fd81e2d-7668-4d24-bc95-5e4ec3d86d5b · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d103bf7-035d-42bb-9beb-54c4f55b5afa · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models An analysis of the use of qualifications on the amazon mechanical turk online labor market
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5e32e358-db83-4460-8e6b-db9fff747a62 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models LAST: Language Model Aware Speech Tokenization
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b94787c0-2378-45a2-b9a6-351247b3a200 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4e6ed73-6450-4012-9043-9e684426521a · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20da324c-57e1-408d-81d9-26b3b92de506 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ccb4176-9e30-49ea-9647-bda5a5e460e9 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models CLAPSpeech: Learning Prosody from Text Context with Contrastive Language-Audio Pre-training
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f66ab1d4-1d32-4423-afed-fe1d397bd964 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Soundstream: An end-to-end neural audio codec
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe38e239-0e49-4449-b1f2-25255f4a6132 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1f7c4c5-f1e1-4549-af60-efd7686a3e4c · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Scaling Speech-Text Pre-training with Synthetic Interleaved Data
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b3a6cc5-aa00-47a4-b7af-0fbc3876389c · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be3cf178-6cc1-4782-80f5-95f27d530a18 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5429ead6-1b82-4c58-b02b-1dea3c8f39f5 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models @esa (Ref
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c77e3bb-9bf3-4d61-80c0-397382e4feb9 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models Unresolved cited work
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db04fd29-6e89-4cdd-9c95-6c6bac061f78 · outbound
ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models One key desirable capability for speech language models is the ability to capture the intricate interdependency between content and prosody
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f56f709b-b828-435b-b0ee-482180d02993 · inbound
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 73a32923-2621-4c7c-9c73-e2c763ebfd53 · inbound
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.