Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-01T01:01:24.536821Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2606.30944.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-01T01:01:24.536821Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 300eb9ff-6fef-4fb1-b473-b40dd573e2f9 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Recent advances in speech language models: A survey
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5fed312d-0459-4563-8504-e014a0f908ed · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation On The Landscape of Spoken Language Models: A Comprehensive Survey
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ff8859a0-4e56-439c-bd81-20434726eb29 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Whislu: End-to-end spoken language under- standing with whisper,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0ea9fb9d-5818-477d-8733-3062002070c3 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Adapting large language model with speech for fully formatted end- to-end speech recognition,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 389a6a33-c15d-43d9-8cf5-abd64ccdf160 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Speechgpt: Empowering large language models with intrinsic cross- modal conversational abilities
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5e81ae1d-a6f6-4baa-9178-f24eeb577039 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Qwen2-Audio Technical Report
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0e0b2984-01f7-4fab-b5c3-ffcd1abad8e1 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 93fe1a6b-ad6a-4f55-ad08-6aff920b25bc · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Understanding the modality gap: An empirical study on the speech-text alignment mechanism of large speech language models,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c4d0555f-3f17-48b2-871e-541e176a93f1 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Alignformer: Modality matching can achieve better zero-shot instruction-following speech-llm,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f041ffa6-11a8-425f-af65-d53aae0d00bb · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Closing the gap between text and speech under- standing in llms
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 902335f4-2250-40f8-80a1-404cd4776d29 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Speech discrete tokens or continuous features? a comparative analysis for spoken language understanding in speechllms,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2f5244f6-835b-4b1b-aa46-d3a482c1ba55 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Closing the Modality Reasoning Gap for Speech Large Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 36f65776-7fda-4dcc-8e86-e7f4e60b5a83 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a5a1814d-23ba-44b9-b394-ba3af7ed3611 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Neural codec language models are zero-shot text to speech synthesizers
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c9232717-a8ce-44d3-b864-9e668f4e3b41 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4ba1e5bb-777c-4576-960b-85610a148913 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b5b13e28-7619-4746-90bf-68e753d28afd · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Pseudo-autoregressive neural codec language models for efficient zero-shot text-to-speech synthesis,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d599cbf4-104d-49dd-baff-0f6bad2c5390 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Indextts2: A breakthrough in emotionally expressive and duration- controlled auto-regressive zero-shot text-to-speech,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 38a83d7b-9754-4a12-a95c-4ce473a84579 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Fish audio s2 technical report
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7fa710d7-b84e-4d7a-a5e8-d77392b18b4e · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Zipvoice-dialog: Non-autoregressive spoken dialogue generation with flow matching,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bb39d317-dfdd-4d0c-91ab-35088b251255 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Moshi: a speech-text foundation model for real-time dialogue
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3cbc1c14-2319-4db7-a69a-b13171bb03a9 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e39f0f1d-08dc-41f0-9592-42e850563044 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Mimo-audio: Audio language models are few-shot learners
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4b757680-88ed-4970-a735-e69877b02bfa · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Kimi-Audio Technical Report
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f667b127-47d2-4a9e-9010-6a67b987b411 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Slam-omni: Timbre-controllable voice interaction system with single-stage training,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6284bc4d-a100-4ddd-a54e-bcf61a6de7ff · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Fun-audio-chat technical report
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 72f844c2-04dc-4cec-9949-df97b3a813a5 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation MOSS-Audio Technical Report
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dd8dc385-a372-410b-8289-a590a05290c1 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b9d06f2e-9fcc-459b-87e8-153712bfeec4 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1147bd70-7f4b-4557-a576-e455ba885218 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Step-Audio 2 Technical Report
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5a8cc085-5f67-4bc4-a36c-60adafae2633 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Qwen2.5-Omni Technical Report
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0f5caf8f-c1c8-4e9e-b8fe-ac27a14719bb · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Qwen3-Omni Technical Report
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 32763740-303e-4d41-870f-d964311c304e · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 152cf258-b15f-43c8-81c8-60d6d204e8b2 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Qwen3.5-Omni Technical Report
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ad9c68d7-5822-4c77-bd6e-dadc3a363a96 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation DeepSeek-V3 Technical Report
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a14c16ae-f3aa-48e1-bf3e-fc1ca63c2bde · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Better & faster large language models via multi-token prediction,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 646895cf-01a7-4a82-aa94-c4542d973006 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Vita-audio: Fast interleaved cross-modal to- ken generation for efficient large speech-language model
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ffef7a5f-8d1f-4a94-8c56-2abd76e49ffb · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation V ocalnet: Speech llm with multi-token prediction for faster and high- quality generation,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 284d6510-4b87-4449-8a25-cf50737a712c · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Similarity of neural network representations revisited
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 54f3c543-76f8-4265-9022-990960a64bf8 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Libriheavy: a 50,000 hours asr corpus with punctuation casing and context,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0a63bf1d-d401-45f8-b874-0cafc1f2d112 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Covost 2 and massively multilingual speech translation,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2a62a2a5-5f38-43e4-a3cf-359a9620df83 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f53f3c67-8aba-40e1-9f0a-f97c394ae11f · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Towards efficient speech-text jointly decoding within one speech language model,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d4fec4bf-a7f2-43ad-9246-5f7d04420371 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Cvss corpus and massively multilingual speech-to-speech translation,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4ef932e0-924c-4dc8-a955-0ffc0f0082a8 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation SLM-S2ST: A multimodal language model for direct speech- to-speech translation,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ac4e1bae-a699-4041-a37b-3006168c7580 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7b7ddd3d-8416-466c-891b-aad929053b09 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Hello gpt-4o,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f7551b2a-e575-4830-96ee-99ca274072c9 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Fleurs: Few-shot learning evaluation of universal representations of speech
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4352eb73-d13b-46e8-9f44-9d7a586a73fb · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Ultraeval-audio: A unified framework for comprehensive evaluation of audio foundation models,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a4ad798f-ee07-4c15-9f5d-1f73655c1a4f · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dbdb48ea-2166-4f36-8e38-5cf7a9447902 · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Vocalbench: Benchmarking the vocal conversational abilities for speech interaction models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation da0035ce-93ca-427e-bcaa-12191012280f · outbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Robust speech recognition via large-scale weak supervi- sion,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
No inbound Pith citation observations are available.