Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T03:53:47.396742Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 93 inbound Pith citation observations for arXiv:2412.02612.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-16T03:53:47.396742Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:33:55.283870Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
50 of 50 outbound references displayed
External citation measurements
2
pith, observed 2026-08-05T02:28:24.338817Z
Observation 428a8c57-b510-4083-ae44-1c68b454d4d4 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3235a4e1-29c7-49e1-9881-79cfcf4a657b · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e3724882-3b3d-4e39-8093-f4b6b5756589 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 81a8ed31-c0a8-4d5b-86d8-0987627100ba · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Tyers, and Gregor Weber
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3fa3a989-7e68-46d1-a0ac-5533eb11bcb2 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Semantic parsing on freebase from question-answer pairs
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a8b21b18-e969-4692-8b1e-c6a1c859dbd0 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Audiolm: A language modeling approach to audio generation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ae4c66a3-9442-4bd7-b22b-2ae082be49b6 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot AISHELL-1: an open-source mandarin speech corpus and a speech recognition baseline
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 78d31843-c909-4228-a158-d1a18b135ad2 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Gigaspeech: An evolving, multi-domain ASR corpus with 10, 000 hours of transcribed audio
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3a63ed3e-23e6-4260-ae60-d5e5710d3ddb · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot SpeechNet: A Universal Modularized Model for Speech Processing Tasks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation af400f69-26f3-42aa-8581-4de09686fd5e · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 128b0c8f-5538-4169-99ad-61eedf01cb26 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot w2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ce7e9348-ed2d-4180-9b74-e1ef2a8d67f6 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot High fidelity neural audio compression
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2b23fb98-f869-48f0-a7b0-ee6330a711b1 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Moshi: a speech-text foundation model for real-time dialogue
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e61d7046-c9c2-4ae4-b5c1-a8dec4d9762d · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Jukebox: A Generative Model for Music
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9624689d-34a5-4bbf-9958-5d558347c2b8 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f6d2f7bb-0212-41b2-821f-8761c3a01feb · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fc3eaf2f-8d26-4f33-b667-85b910de8067 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1a4a10bf-1b01-4729-90a6-2af436a5ba2b · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Textually pretrained speech language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 71edbb6f-c061-453f-9cd6-c1f330ff99ee · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Visqol: an objective speech quality model
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3a9dac0b-07b1-4167-844f-875b90cd86b7 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e2b763b3-d87f-49df-8695-4eb8f22c3ee3 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b0890416-7408-4d34-a069-6cdd6a38cf59 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Weld, and Luke Zettlemoyer
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b2135a48-3780-4905-b892-b91538d23279 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ecca975d-faad-4b7f-9616-7c05c0bbaa02 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot High-fidelity audio compression with improved RVQGAN
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 34f5ceb6-10a5-41d6-864e-315132f2e335 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot On generative spoken language modeling from raw audio
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ada80ec3-6120-48fb-a3c9-5fc21ccabf5f · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Hashimoto
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation eef3f15d-3231-419d-b0c4-8ca8a977832f · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Mosnet: Deep learning-based objective assessment for voice conversion
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3f6fb731-0617-4ffb-9282-35d5f64b2ea1 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Decoupled Weight Decay Regularization
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 946fcd8e-2a4f-4fb0-a601-248d42e628ab · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Matcha-TTS: A fast TTS architecture with conditional flow matching
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cf78982d-c7b2-469a-882c-50d91a7f073d · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot A Corpus and Evaluation Framework for Deeper Understanding of Commonsense Stories
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 819280b5-76a1-44e4-b111-a8954543a1a2 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8f72d4a3-3105-4f5e-98b6-9b7137c48b87 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Expresso: A benchmark and analysis of discrete expressive speech resynthesis
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c4fbebb8-c809-42ba-9429-97f5c5813b31 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Spirit LM: Interleaved Spoken and Written Language Model
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 56f4ccd2-1fab-4072-959a-883767a84c87 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Hello gpt-4o
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation da8ccf8d-a09a-43c0-883e-8ab295c48875 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Librispeech: An asr corpus based on public domain audio books
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7a9de88f-0a4d-48aa-8860-e228c2bdc72c · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot MLS: A large-scale multilingual dataset for speech research
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c7df4242-ba72-4665-bea6-3ed93ee2f1a0 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Robust speech recognition via large-scale weak supervision
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 27d591dd-3b16-4bd0-9f57-488deefa01fc · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Utmos: Utokyo-sarulab system for voicemos challenge 2022
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 93ab99dc-015a-47a7-ad76-78ab99bfb764 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot SeACo-Paraformer: A Non-Autoregressive ASR System with Flexible and Effective Hotword Customization Ability
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3d330d51-644f-45ba-8f7d-b6873c7187df · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Senior, and Koray Kavukcuoglu
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fe09de31-23ef-4236-acc5-6a6450230f34 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Neural discrete representation learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation abae069a-dbed-4db4-87c3-2b4187cdda19 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 65b8e58d-971c-46f6-918a-71daacd9e47c · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f4700e4e-0ad4-49df-a1f1-30827606ba70 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Wenet: Production oriented streaming and non-streaming end-to-end speech recognition toolkit
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2baaaea2-be98-4cc4-9ba7-08c715ce5c63 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Soundstream: An end-to-end neural audio codec
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 140afabd-8852-45ff-b4fd-3259ed7362f6 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Scaling Speech-Text Pre-training with Synthetic Interleaved Data
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ef1432ba-bd84-420b-b16e-d9df22ac1ef1 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f8751651-4312-4aab-aba1-28dbbaf552d9 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot OPT: Open Pre-trained Transformer Language Models
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2c157565-0887-4b56-952c-48550c350d74 · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Speechtokenizer: Unified speech tokenizer for speech language models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 13dd085b-36e2-49f1-b5c6-9e8b5f9fdc4b · outbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9f82f6b4-caf6-49c0-9063-ba3a842259f1 · inbound
InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 137
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81f76123-dd3a-464e-91bc-16ec47d7dc7c · inbound
SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c2aa9ce-c2a9-4cdd-b070-202333293215 · inbound
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 365b9f81-3b01-4293-9079-df43b2aeaeaa · inbound
Real-Time Textless Dialogue Generation GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d5e7723-6d36-414f-a1d0-1646215d8e72 · inbound
VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82f5b4c4-359b-4b56-a715-6311a995534e · inbound
MinMo: A Multimodal Large Language Model for Seamless Voice Interaction GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9e5eeee-9948-49f2-91ae-23aaaca74e2e · inbound
LUCY: Linguistic Understanding and Control Yielding Early Stage of Her GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6ccf608-f3df-4855-b006-c7740879e3e6 · inbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e4f0b3ff-3453-4591-a366-942211b6c35c · inbound
Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b3330311-5418-4f13-bf52-3f0ff7beb550 · inbound
On The Landscape of Spoken Language Models: A Comprehensive Survey GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e7837def-94de-4770-8752-6fe5906d3df1 · inbound
Kimi-Audio Technical Report GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 72888da8-50f9-4add-9bb4-6f751b1f8573 · inbound
S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39200706-e687-4c09-9335-2c5b907af55e · inbound
SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a92c2aa1-99ea-4f6b-8ef0-909ed81a58a6 · inbound
Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d70c55fd-bf15-4950-beb6-956a139fbbc8 · inbound
MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d11e388d-38c6-4330-946b-5ce5a959d5d2 · inbound
OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b15efbd-c7c1-496d-b92e-9470a0a2afa2 · inbound
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ad31277-d93e-4068-a1cd-f9b30221ff29 · inbound
TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74adbc61-b926-4cf6-9355-eb9286bd6796 · inbound
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6364a05c-56a0-4e4d-a795-8ba0b4bfed1d · inbound
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 171ac415-b9c8-47de-aab3-1bd0b16a5aec · inbound
Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27cd1f78-1374-4e2a-a5cd-7dad6adb1e06 · inbound
Debunk and Infer: Multimodal Fake News Detection via Diffusion-Generated Evidence and LLM Reasoning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ed8cb77-4e20-4711-b181-e625ac32b5ec · inbound
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e1311b8-b12f-4de8-af07-d1dff9397c44 · inbound
StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1b2a798-1d33-450d-853d-7e9925e973bc · inbound
Step-Audio 2 Technical Report GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 134ce70d-1692-43cd-9080-f731de608cff · inbound
BoSS: Beyond-Semantic Speech GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73677d9d-6346-4a6e-81e2-7ba5cb3d7b94 · inbound
LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 494656bc-0d91-41ad-b14b-d49cb69a3022 · inbound
Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c5927f2-6892-4b69-8a73-7dd5735b9669 · inbound
VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4407fe5-fb24-4538-b95f-ae677ddbc10a · inbound
An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b720b867-8b51-4b57-821a-a6cd1d1ae847 · inbound
Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 98b2d878-c61b-4ef6-bea8-a1f6228f8976 · inbound
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6d8d4062-089e-447a-aa45-d6522747308d · inbound
AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 51cab11e-1597-4bca-b4bb-ce67cf821c5d · inbound
Make a Video Call with LLM: A Measurement Campaign over Six Mainstream Apps GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d8c3881-632a-40ca-9478-ee325d43010f · inbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2f14981-1b32-4f39-a32d-6c17d8364115 · inbound
Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ae9f91cc-8f7b-4b89-9e61-e0631934f4d4 · inbound
VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f163327-00b8-4f3a-b94d-281814f3111c · inbound
ORCA: Open-ended Response Correctness Assessment for Audio Question Answering GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75ad7880-3258-4cda-8406-36d20049e101 · inbound
ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 129
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e06ff900-f494-4be5-aec2-f0b60a17301b · inbound
Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ff6404e0-db9d-4ad3-9595-3a61f2f7a93e · inbound
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24d8b381-7d8d-4761-8d9d-7a3e93f2bca4 · inbound
Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aecb237b-c5fb-4829-848d-db0df7e067c6 · inbound
The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4edbecd6-0c16-47a1-9d78-9480c99cf225 · inbound
The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b84ee168-78db-49ea-a7d2-ab11bedccfc5 · inbound
The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cdda8ef-ac8a-4653-9721-18588a276599 · inbound
TiCo: Time-Controllable Spoken Dialogue Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5945c65c-a5a3-41ba-ae7d-78afb580d508 · inbound
Sharp spectral estimates for free boundary problems arising in plasma physics GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e02c3ed-03eb-4df1-85c8-3945770da139 · inbound
FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f79f6af8-d026-4f35-a0cd-278bcafbacf1 · inbound
Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52e90ac9-568a-4373-bad9-cc5b2a1fb956 · inbound
Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ef66c270-e501-472a-aed7-24e6d05efed4 · inbound
GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e414a280-1a0a-4bde-8880-b714bf72858e · inbound
HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d0a4d9f3-2bb2-4bd5-bb3e-8e7585d86f17 · inbound
Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4e1f9513-42b1-4389-886f-170370dafe4b · inbound
A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3332b3a4-1a5c-4135-a0ad-e8c834f7deab · inbound
Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b1ee6800-90a8-4286-8196-c135310f7c5b · inbound
Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4b3b2aca-f622-42ec-a386-c53388929e1b · inbound
SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 78159a47-e9c3-4a35-b3f4-c571a44aa3bb · inbound
Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a3e5a554-578a-4010-801a-8c6f61e5d8d6 · inbound
Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 088b1b3b-a760-4c2a-9f1c-c1e434bbbf05 · inbound
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ae6a8181-0131-4116-baea-9a6d8f374082 · inbound
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c4aa453e-1067-4b02-8b1d-d7ea0b94f867 · inbound
VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b15eccef-3678-42c8-a3c9-70e638da3ff3 · inbound
How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 98d73a50-35cc-412a-ad0d-d42c203f2960 · inbound
AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2e492102-95ed-4ce4-a186-3f55b82241bc · inbound
A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0c127363-d584-4c42-a74d-b5d003087eb3 · inbound
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 85596773-8041-411c-8de9-8380400cefa5 · inbound
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 16a25c88-bcf3-4bab-9a14-93d06c2c3f09 · inbound
A Survey of Audio Reasoning in Multimodal Foundation Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation edc17ada-56ba-49c9-86e3-10d9c1489f08 · inbound
Toward Native Multimodal Modeling: A Roadmap GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cbba8d8f-bc45-413a-9fd7-63d48dc29d96 · inbound
Learning When to Think While Listening in Large Audio-Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c75e5f79-e612-424b-ab5e-08c82c1d9955 · inbound
LaSR: Context-Aware Speech Recognition via Latent Reasoning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ddb2aaee-48dd-43de-8915-fc12e7d63ee2 · inbound
Sympatheia: Emotionally Adaptive Voice Assistant with Continuous Affect Conditioning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 363b0d7f-2e34-4f19-b7e8-1c3d336660ec · inbound
MOSS-Audio Technical Report GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 46d7b5d9-f76d-49b4-a521-2b4d25093144 · inbound
IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 14c18da9-d85b-446c-a9f6-ca86dbe854ac · inbound
Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 51b0275c-260b-4804-a63b-b55679767853 · inbound
Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2f7da106-8fab-442b-9c8f-1b5a96b9c611 · inbound
Benchmarking Neural Speech Compression from a Rate-Distortion Perspective GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ed2af017-186b-431c-babe-07a59fa63cfe · inbound
Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2f4fe339-2ce1-4cc4-baa6-127bc09eb08f · inbound
Endpoint Anticipation for Low-Latency Spoken Dialogue GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ba53d677-58c0-49fc-a6a8-0252b2c5af32 · inbound
Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dd8dc385-a372-410b-8289-a590a05290c1 · inbound
Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e2d1cb7f-9588-4af1-8b32-c1d1f461a2c8 · inbound
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 212
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bb37c311-8601-4b7d-a64f-07654ece0ec2 · inbound
Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 300a1252-3d10-4733-87c3-75c4314afba4 · inbound
Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ee5c60a1-fc3e-4c11-939c-c96a57f55662 · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 175
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5b79b252-f52f-44bd-b2a5-259558dc1f95 · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 175
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2226ca2-91bf-4957-b01f-fb74290ba53e · inbound
Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 99fb2f1a-4f93-43a9-8bd2-85cdf477a05a · inbound
Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebb9399c-db7c-4ad2-a17d-9402744aee6b · inbound
Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 798d44f9-d625-476e-b436-eb0ef3435f70 · inbound
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10adce2e-e654-40fc-9766-7909b25b3abd · inbound
SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c648a9b2-30be-496d-9dd7-ddc0d543d6e7 · inbound
DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9515d57e-b62e-4052-85c9-8c62812abc88 · inbound
EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.