Pith. sign in

Paper Citation Record · LEDGER

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

As of 12 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 93 inbound Pith citation observations for arXiv:2412.02612.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02612 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T03:53:47.396742Z

measured 143 of 143 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 93 of 93 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:33:55.283870Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact20
  • verified fuzzy27
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 428a8c57-b510-4083-ae44-1c68b454d4d4 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.534947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:422871d152dab8ed1e8a4e70d7abd05290dd8d1ebdb401c56009b084411ff5d7

Observation 3235a4e1-29c7-49e1-9881-79cfcf4a657b · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.451410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:cc82c897189d2c786fb899638b9871c2261335fa6627e25c5b384ce1a517c7da

Observation e3724882-3b3d-4e39-8093-f4b6b5756589 · outbound

This paper cites Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.618077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:1e7fd92c8afea604e7a1562e22cc99ba32d4869aad50f0201816711b7e0a375a

Observation 81a8ed31-c0a8-4d5b-86d8-0987627100ba · outbound

This paper cites Tyers, and Gregor Weber.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Tyers, and Gregor Weber

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.622424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:dee744ea735f40b9acc7bc2fb9d036062fcb93b30b7a83e4cb69d58442e1b4c4

Observation 3fa3a989-7e68-46d1-a0ac-5533eb11bcb2 · outbound

This paper cites Semantic parsing on freebase from question-answer pairs.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Semantic parsing on freebase from question-answer pairs

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.626624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:7b0147f553772ea26332493eb9eeaa018b33edb1ea5768ad7ad89e2d8b7b52fd

Observation a8b21b18-e969-4692-8b1e-c6a1c859dbd0 · outbound

This paper cites Audiolm: A language modeling approach to audio generation.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Audiolm: A language modeling approach to audio generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.630438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:8bf8747a4271a09ebf1239d59cf135ceeee5209445556214b7edf849dc507fd9

Observation ae4c66a3-9442-4bd7-b22b-2ae082be49b6 · outbound

This paper cites AISHELL-1: an open-source mandarin speech corpus and a speech recognition baseline.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot AISHELL-1: an open-source mandarin speech corpus and a speech recognition baseline

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.634892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:cb75c224ce3e161bc86c4f714f8cc720aa22e52c76819a7c80682e29b43c1ab8

Observation 78d31843-c909-4228-a158-d1a18b135ad2 · outbound

This paper cites Gigaspeech: An evolving, multi-domain ASR corpus with 10, 000 hours of transcribed audio.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Gigaspeech: An evolving, multi-domain ASR corpus with 10, 000 hours of transcribed audio

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.639545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:4dc85a3c2279421c4158a9d9c26e2667e773495415a73ddd912f92fc079172f1

Observation 3a63ed3e-23e6-4260-ae60-d5e5710d3ddb · outbound

This paper cites SpeechNet: A Universal Modularized Model for Speech Processing Tasks.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot SpeechNet: A Universal Modularized Model for Speech Processing Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.568840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:587096a1d270c8f041950d984f96051cfe1f3e7ed25ebeab4cba4f4daef1598f

Observation af400f69-26f3-42aa-8581-4de09686fd5e · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:53:47.586781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:d3d537373ce1c7284ae41317ac3648ead165530bbee0c8a7846c9b79903966af

Observation 128b0c8f-5538-4169-99ad-61eedf01cb26 · outbound

This paper cites w2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot w2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.643304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:cf32cc8ec841aefe6fab1923d1bb39c6199555d23f7d48b302a1fa17a6b8f26c

Observation ce7e9348-ed2d-4180-9b74-e1ef2a8d67f6 · outbound

This paper cites High fidelity neural audio compression.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot High fidelity neural audio compression

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.647053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:b3921d34bffe1290ce855d08559845140b70fe6c3493b9fde0cea6a4ba5d1081

Observation 2b23fb98-f869-48f0-a7b0-ee6330a711b1 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Moshi: a speech-text foundation model for real-time dialogue

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.651713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:9a375bbd5010047c6edc99d561e1a2282a9acc0c67d1844a420b0bf59316e6e3

Observation e61d7046-c9c2-4ae4-b5c1-a8dec4d9762d · outbound

This paper cites Jukebox: A Generative Model for Music.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Jukebox: A Generative Model for Music

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:53:47.474084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:3f7535455a6c1dae854aa00f6fa7cac449b7ebb13f103bee91ae0d5e1370c8a3

Observation 9624689d-34a5-4bbf-9958-5d558347c2b8 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:53:47.489342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:c256fcae56a3ba224d9d5fb111c229599b43b2b7d37ce8f38a175b8ed1d46683

Observation f6d2f7bb-0212-41b2-821f-8761c3a01feb · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.496972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:a8f995eba70bfd18c2558fc5aef8a687d3b45b9bd7e4827da3ade18297a1eb9f

Observation fc3eaf2f-8d26-4f33-b667-85b910de8067 · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:53:47.504259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:12eb9e229a81750d411667005e8b62ca245a742bbc0b8fa347fae9153104b41e

Observation 1a4a10bf-1b01-4729-90a6-2af436a5ba2b · outbound

This paper cites Textually pretrained speech language models.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Textually pretrained speech language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.655384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:de64d6412bed2f16cb88c1469398154443405835db58bca8f6483b51f18888a4

Observation 71edbb6f-c061-453f-9cd6-c1f330ff99ee · outbound

This paper cites Visqol: an objective speech quality model.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Visqol: an objective speech quality model

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.658897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:c0a798c2bf8092e519864c4fe83b244a3d1b33d06d8c867eea4e58626731ded6

Observation 3a9dac0b-07b1-4167-844f-875b90cd86b7 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.662743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:bc2ec8675daca8ee7f6281e7066c811e846ee62e2dedca500ffc663f32a0a814

Observation e2b763b3-d87f-49df-8695-4eb8f22c3ee3 · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.459620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:56c7f0b77fdff1ffe1ad309cee75690dbdd6d3b2a85223fd676e1a87a115544c

Observation b0890416-7408-4d34-a069-6cdd6a38cf59 · outbound

This paper cites Weld, and Luke Zettlemoyer.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Weld, and Luke Zettlemoyer

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.666570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:b353c040b07d181ab70dfd92b5b7d558ae20ff0d150c7d53ed753168ce0a5cf6

Observation b2135a48-3780-4905-b892-b91538d23279 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.670373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:76587141f2fffa09570611d648ceae7967baed4205e054b71d7498c01ab5248d

Observation ecca975d-faad-4b7f-9616-7c05c0bbaa02 · outbound

This paper cites High-fidelity audio compression with improved RVQGAN.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot High-fidelity audio compression with improved RVQGAN

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.674339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:3f1d3ff9b4e0fefe1f27de721db296b586fe446452928a19acfdb9fe2cb689bf

Observation 34f5ceb6-10a5-41d6-864e-315132f2e335 · outbound

This paper cites On generative spoken language modeling from raw audio.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot On generative spoken language modeling from raw audio

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.678369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:521607d6eccf6cdf05a823378def925e441ce5e6152e8ef7aab5edcbec136659

Observation ada80ec3-6120-48fb-a3c9-5fc21ccabf5f · outbound

This paper cites Hashimoto.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Hashimoto

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.681842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:5e6f8db1a1a190b4972a985a61a8f5d1b1bb6eb0b4487c24cb733a90e8cb58cb

Observation eef3f15d-3231-419d-b0c4-8ca8a977832f · outbound

This paper cites Mosnet: Deep learning-based objective assessment for voice conversion.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Mosnet: Deep learning-based objective assessment for voice conversion

Reference 27

Resolution
verified exact
doi, observed 2026-05-16T03:53:47.443984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:8d62a5e161a5bbbffec8941519de8c66435de67e3c408f7d4ea79dd1b9d33e40

Observation 3f6fb731-0617-4ffb-9282-35d5f64b2ea1 · outbound

This paper cites Decoupled Weight Decay Regularization.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Decoupled Weight Decay Regularization

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:53:47.561835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:69a44c63c1fe5e686f888d2628f31318276daebd798ce8fa9b898230ea062c12

Observation 946fcd8e-2a4f-4fb0-a601-248d42e628ab · outbound

This paper cites Matcha-TTS: A fast TTS architecture with conditional flow matching.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Matcha-TTS: A fast TTS architecture with conditional flow matching

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.685465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:b0437b36c6da91023b19784b73f04f48ed3e47c828d41bcf4a3bdb2128ce0941

Observation cf78982d-c7b2-469a-882c-50d91a7f073d · outbound

This paper cites A Corpus and Evaluation Framework for Deeper Understanding of Commonsense Stories.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot A Corpus and Evaluation Framework for Deeper Understanding of Commonsense Stories

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:53:47.580477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:dca015678fdc11d2905342e91441bdad17310e5f6c0fe23ea69520415d3c2fbd

Observation 819280b5-76a1-44e4-b111-a8954543a1a2 · outbound

This paper cites an unresolved cited work.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-16T03:53:47.689123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:b5c45a133158d34f60642ccfce415a8c9622ae54bb717488cb7f630eceb29ea1

Observation 8f72d4a3-3105-4f5e-98b6-9b7137c48b87 · outbound

This paper cites Expresso: A benchmark and analysis of discrete expressive speech resynthesis.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Expresso: A benchmark and analysis of discrete expressive speech resynthesis

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.693009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:204eea0fb3b15df85449897fa69d9e886b08a35d06bfb33faeeaabb2aaf8a3e9

Observation c4fbebb8-c809-42ba-9429-97f5c5813b31 · outbound

This paper cites Spirit LM: Interleaved Spoken and Written Language Model.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Spirit LM: Interleaved Spoken and Written Language Model

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.467367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:948e60f60211dc2768711b55c7323b806f59e08b51c19f44b00f003d5255ab1c

Observation 56f4ccd2-1fab-4072-959a-883767a84c87 · outbound

This paper cites Hello gpt-4o.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Hello gpt-4o

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.696496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:13bdca7cfefe4c7362a9fc67d69a5196c92f36314192161cb63708ad7f1270b9

Observation da8ccf8d-a09a-43c0-883e-8ab295c48875 · outbound

This paper cites Librispeech: An asr corpus based on public domain audio books.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Librispeech: An asr corpus based on public domain audio books

Reference 35

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T03:53:47.482142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:8a409c6fb5780652d12ddcbd43f08c0b540329896cb6a455b55dd4d204ac5dad

Observation 7a9de88f-0a4d-48aa-8860-e228c2bdc72c · outbound

This paper cites MLS: A large-scale multilingual dataset for speech research.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot MLS: A large-scale multilingual dataset for speech research

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.591044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:4a53033e1f4dce73507a2f998a705dfcf86a8167edd763a379c57b2fb1b8a898

Observation c7df4242-ba72-4665-bea6-3ed93ee2f1a0 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Robust speech recognition via large-scale weak supervision

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.595307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:407b3c8302a8dd4d3eca8a8fec87a5f4ffece656ebc761a736534a51b1fdb4ef

Observation 27d591dd-3b16-4bd0-9f57-488deefa01fc · outbound

This paper cites Utmos: Utokyo-sarulab system for voicemos challenge 2022.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Utmos: Utokyo-sarulab system for voicemos challenge 2022

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.598875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:9a582a9bf74e0d6344db6af3f1dca06acd2373138892a79a705bc2363c14028a

Observation 93ab99dc-015a-47a7-ad76-78ab99bfb764 · outbound

This paper cites SeACo-Paraformer: A Non-Autoregressive ASR System with Flexible and Effective Hotword Customization Ability.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot SeACo-Paraformer: A Non-Autoregressive ASR System with Flexible and Effective Hotword Customization Ability

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.512254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:39a70ce5f6063069dbcba14bf06706e0262c09abb6ddf21564f7483027a4bf29

Observation 3d330d51-644f-45ba-8f7d-b6873c7187df · outbound

This paper cites Senior, and Koray Kavukcuoglu.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Senior, and Koray Kavukcuoglu

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.602159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:56f4c08312792083b4084e9d0863861759e71088660bd8c602881cd7ebc37b9e

Observation fe09de31-23ef-4236-acc5-6a6450230f34 · outbound

This paper cites Neural discrete representation learning.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Neural discrete representation learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.606181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:4cea05685b5f647227fa3c952e2b31fc596c165b565e2d85f4f75f2e762ac336

Observation abae069a-dbed-4db4-87c3-2b4187cdda19 · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.520398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:bf6eb44c239469560914319e0f00c4c118de5795f429d6ce0a7f7f213b9e5cf3

Observation 65b8e58d-971c-46f6-918a-71daacd9e47c · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.527359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:db8133cbefb8b2d5488594fc235c007b4c0af83baba5f065f2c6ddba9d4677b7

Observation f4700e4e-0ad4-49df-a1f1-30827606ba70 · outbound

This paper cites Wenet: Production oriented streaming and non-streaming end-to-end speech recognition toolkit.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Wenet: Production oriented streaming and non-streaming end-to-end speech recognition toolkit

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.614235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:3cc579c48f08574a70001a6a30605759456d74503b01d0d8f1a08ea8485e3767

Observation 2baaaea2-be98-4cc4-9ba7-08c715ce5c63 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Soundstream: An end-to-end neural audio codec

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.438327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:2fe5becba0a1bf4f572737039b4f1e3c5007c782e9baab9f2b17bcedbb102cb0

Observation 140afabd-8852-45ff-b4fd-3259ed7362f6 · outbound

This paper cites Scaling Speech-Text Pre-training with Synthetic Interleaved Data.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Scaling Speech-Text Pre-training with Synthetic Interleaved Data

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.541943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:e7844e9a5f44aa4bf22e5b5032bb3c2071cffd5a71629508c6505ba5c4cde660

Observation ef1432ba-bd84-420b-b16e-d9df22ac1ef1 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.549814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:ac5cdcca3850fda8a0a2ff7155d27e56b3e19fa82363b10b647e33622dd0cde9

Observation f8751651-4312-4aab-aba1-28dbbaf552d9 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot OPT: Open Pre-trained Transformer Language Models

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T03:53:47.556162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:b36abe9390a5d85a493e52638bd69e1b0e79604324a3b2ffd7aff3a148aa39ea

Observation 2c157565-0887-4b56-952c-48550c350d74 · outbound

This paper cites Speechtokenizer: Unified speech tokenizer for speech language models.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Speechtokenizer: Unified speech tokenizer for speech language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T03:53:47.610346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:662032ddacc27d9c043f78d8e4e50f0b341d7e3ec49bf7d114494c6d0afcc965

Observation 13dd085b-36e2-49f1-b5c6-9e8b5f9fdc4b · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-16T03:53:47.574522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:27f947120d2a4714aeb533e8eeb7ee48f93aa1b118578f2c757661b3cf49f563

Pith citing papers

Observation 9f82f6b4-caf6-49c0-9063-ba3a842259f1 · inbound

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions cites this paper.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 137

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.551972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.551972Z digest=sha256:fe6cd28d1118cc7b6a22d5891cf3ec6334a7f4cbb0c6ae7f26d14b5f3d16ffeb

Observation 81f76123-dd3a-464e-91bc-16ec47d7dc7c · inbound

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training cites this paper.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.459612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.459612Z digest=sha256:31573170b94652d33f2677900bf354cecb8cfeba08037767e3ba562d07a59ff0

Observation 1c2aa9ce-c2a9-4cdd-b070-202333293215 · inbound

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis cites this paper.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:49.064735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:49.064735Z digest=sha256:88b2ca01b128f35055dd6b1e55ef46bb66c9544bcf59ad1d9b56065d809c2d91

Observation 365b9f81-3b01-4293-9079-df43b2aeaeaa · inbound

Real-Time Textless Dialogue Generation cites this paper.

Real-Time Textless Dialogue Generation GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T21:27:45.655157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:27:45.655157Z digest=sha256:252ea186a82df70d7b605ef40734d8ca3a3bd67eca289f6d0400ea2b749d3aec

Observation 3d5e7723-6d36-414f-a1d0-1646215d8e72 · inbound

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models cites this paper.

VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T21:24:18.278793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:24:18.278793Z digest=sha256:09bb18834c06a0e45d557f01c65377189765c4536600287811c12187dc07e173

Observation 82f5b4c4-359b-4b56-a715-6311a995534e · inbound

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction cites this paper.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.860538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.860538Z digest=sha256:d5a04687f8ce4a909d8ac39e56fb98f53345ebb14b615d1d51336c1daafbb2cd

Observation e9e5eeee-9948-49f2-91ae-23aaaca74e2e · inbound

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her cites this paper.

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T13:34:25.715053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:34:25.715053Z digest=sha256:5c6a4dd8dcf12dc3c5739cf7e8dc33f16c9c18cd855c7e62806454d032a52b6d

Observation b6ccf608-f3df-4855-b006-c7740879e3e6 · inbound

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction cites this paper.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:39:48.403186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:8b6131ba46f208d2a5dc61d0b6e5c1a3d73712109276a725c08c94a81564884f

Observation e4f0b3ff-3453-4591-a366-942211b6c35c · inbound

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs cites this paper.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:54ef487bc84a21e8b6f7b2168eca4259586623fecd8b4ea83985f2e4578ecef7

Observation b3330311-5418-4f13-bf52-3f0ff7beb550 · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-22T20:45:08.172164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:b6d4fcb29285d730edcf7b834143fd1c317576b25a3f4f0b36809c7206300bd9

Observation e7837def-94de-4770-8752-6fe5906d3df1 · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:45171359ae81b90ac0dbcbc52a97579786b3c9c0fbc6723df701c7ec1b9b4ee0

Observation 72888da8-50f9-4add-9bb4-6f751b1f8573 · inbound

S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models cites this paper.

S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:38.037256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:38.037256Z digest=sha256:e46e0f78de349c2f0fdf7d42297bc6dd1f7ef7c36a78ebdb1ee1c612af4e5633

Observation 39200706-e687-4c09-9335-2c5b907af55e · inbound

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model cites this paper.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.059351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.059351Z digest=sha256:570812259d601bc6e10fa972ab5363251a047ef03e34fef24bf2d269b3349770

Observation a92c2aa1-99ea-4f6b-8ef0-909ed81a58a6 · inbound

Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models cites this paper.

Analyzing Mitigation Strategies for Catastrophic Forgetting in End-to-End Training of Spoken Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:49:29.942638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:49:29.942638Z digest=sha256:96a4b6b7d459d8e8f0add21570d0f016be75a4688a7aa781b91b87f4924ff7fe

Observation d70c55fd-bf15-4950-beb6-956a139fbbc8 · inbound

MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding cites this paper.

MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:33.837927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:33.837927Z digest=sha256:a13df7f32a0e0f74e7e690e0589c4740ee9cebaff2172084fbf10b95a22ac9ca

Observation d11e388d-38c6-4330-946b-5ce5a959d5d2 · inbound

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction cites this paper.

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:33.352219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:01:33.352219Z digest=sha256:a19db7220598b5a4db81857091bea2883de843255fec19ec6db063efc0b7e515

Observation 8b15efbd-c7c1-496d-b92e-9470a0a2afa2 · inbound

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant cites this paper.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.320594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.320594Z digest=sha256:d0d17de34e462bbe6272f305635f795985f244482a637658aa59dbd317d6d2c8

Observation 6ad31277-d93e-4068-a1cd-f9b30221ff29 · inbound

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment cites this paper.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:42.585526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:42.585526Z digest=sha256:93f048bc4b305b58fbb8969952ead0972fb35554add51943d4e539592ecaf7ff

Observation 74adbc61-b926-4cf6-9355-eb9286bd6796 · inbound

In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion cites this paper.

In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:50:27.902521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:50:27.902521Z digest=sha256:a881febe45e7a8ba3cc75bc62328fc40a33e7f9274136b7ee76b1e3171863f58

Observation 6364a05c-56a0-4e4d-a795-8ba0b4bfed1d · inbound

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model cites this paper.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:15.913730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:15.913730Z digest=sha256:3aa5a5d914fbf243c6b336605fc57202b4cf66363ab6c4d5583d307d84e8b49b

Observation 171ac415-b9c8-47de-aab3-1bd0b16a5aec · inbound

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model cites this paper.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.263029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.263029Z digest=sha256:3ca238da9a191a41553a2043ef05d3e23fa5deaaf398c33161bcf6cd4fda3254

Observation 27cd1f78-1374-4e2a-a5cd-7dad6adb1e06 · inbound

Debunk and Infer: Multimodal Fake News Detection via Diffusion-Generated Evidence and LLM Reasoning cites this paper.

Debunk and Infer: Multimodal Fake News Detection via Diffusion-Generated Evidence and LLM Reasoning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:51:45.320447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:51:45.320447Z digest=sha256:5995f7706d2300e0af5fc5a06ef6409df90db3a747e7b6a40f00c965d2b824c0

Observation 7ed8cb77-4e20-4711-b181-e625ac32b5ec · inbound

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs cites this paper.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:36.668481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:36.668481Z digest=sha256:1926a71feda8db369b5d5b4850f3091954873101656e2d8734474e3ee05188ca

Observation 4e1311b8-b12f-4de8-af07-d1dff9397c44 · inbound

StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding cites this paper.

StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:33:09.263696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:33:09.263696Z digest=sha256:9f4139fdbe847648e7ee34b3b642bd1e175d1ad14ae692fbc8879af276658ec4

Observation c1b2a798-1d33-450d-853d-7e9925e973bc · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-16T05:59:51.104219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:3c745a8975ae2caa37bf3dc24ec08ff9d74796a4253870d3d861b22ad70381ba

Observation 134ce70d-1692-43cd-9080-f731de608cff · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.294343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.294343Z digest=sha256:aafa116c6fd49d42f3a1d2eb5c9fe49c1785934e7c5d96831363fa9399d7bd52

Observation 73677d9d-6346-4a6e-81e2-7ba5cb3d7b94 · inbound

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model cites this paper.

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-05T17:56:56.634814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:56:56.634814Z digest=sha256:0049a9c23c1b91b04ee26d6b677877d91b27d62b2498619d93be0ad6d449e1be

Observation 494656bc-0d91-41ad-b14b-d49cb69a3022 · inbound

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs cites this paper.

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T16:46:49.270919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:46:49.270919Z digest=sha256:fb104280d7a9e100c44036b08ec57ae3d5c604269d3eea1c69d67fda475da348

Observation 8c5927f2-6892-4b69-8a73-7dd5735b9669 · inbound

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents cites this paper.

VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:06.477526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:06.477526Z digest=sha256:03dc34402a0247b2adfafd62c311403cbadd11e828e304805b22a892df4e7883

Observation a4407fe5-fb24-4538-b95f-ae677ddbc10a · inbound

An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training cites this paper.

An Empirical Analysis of Discrete Unit Representations in Speech Language Modeling Pre-training GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T10:53:41.369932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:53:41.369932Z digest=sha256:a1f93056f1712af02df71c4c36fa8bdc9663912a3f013b75e3908e8630cf78fd

Observation b720b867-8b51-4b57-821a-a6cd1d1ae847 · inbound

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages cites this paper.

Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-18T16:31:37.251470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T16:27:37.596817Z digest=sha256:6191e858483ec11115fc35b9eb7074b524fc90983357de9d6bfef1c29dd9fd83

Observation 98b2d878-c61b-4ef6-bea8-a1f6228f8976 · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:01:24.330328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:4b38834b8d9387048d144903c29ed9e154eee0a348bddda97ddae7c8ec465119

Observation 6d8d4062-089e-447a-aa45-d6522747308d · inbound

AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models cites this paper.

AudioRole: An Audio Dataset for Character Role-Playing in Large Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:06:23.831872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T13:03:50.876235Z digest=sha256:fe9cb44a4ac62b7b6234aa1befeb466d36d49c970ea9877354780f21b04d7115

Observation 51cab11e-1597-4bca-b4bb-ce67cf821c5d · inbound

Make a Video Call with LLM: A Measurement Campaign over Six Mainstream Apps cites this paper.

Make a Video Call with LLM: A Measurement Campaign over Six Mainstream Apps GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T13:29:04.277748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:29:04.277748Z digest=sha256:4e396ee1cd15bb94999ecd80e34f0a01788cc3180b1f5e3feb4e6d7420d055bc

Observation 6d8c3881-632a-40ca-9478-ee325d43010f · inbound

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues cites this paper.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:44.532704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:44.532704Z digest=sha256:34fffac49b0572f6eb564ec03c31c4b801e840aaf7a84c77c4397a6956278db8

Observation b2f14981-1b32-4f39-a32d-6c17d8364115 · inbound

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models cites this paper.

Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-18T07:46:03.570346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T07:43:23.913399Z digest=sha256:f558e57407c1895150456a834bf61515f11ac810e274fd0ce8578868bc39622f

Observation ae9f91cc-8f7b-4b89-9e61-e0631934f4d4 · inbound

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents cites this paper.

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T10:13:43.130162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:13:43.130162Z digest=sha256:11b16e14ec3dbcbbcdbf7c38ef2b14cd850f6839780f9d68c9ef120e8246c84e

Observation 0f163327-00b8-4f3a-b94d-281814f3111c · inbound

ORCA: Open-ended Response Correctness Assessment for Audio Question Answering cites this paper.

ORCA: Open-ended Response Correctness Assessment for Audio Question Answering GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T19:38:26.212573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T19:38:26.212573Z digest=sha256:c12b58fdc2d5a65d2b6af974ed95b1c153bce2a78d5144f8e8a5477316d1da8c

Observation 75ad7880-3258-4cda-8406-36d20049e101 · inbound

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body cites this paper.

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 129

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:08:36.269780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T22:04:07.403410Z digest=sha256:cda9213d2bda40a6b154a7a4dc9826c313e71e378183741bfd0184587f562867

Observation e06ff900-f494-4be5-aec2-f0b60a17301b · inbound

Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models cites this paper.

Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:28:20.116056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-16T19:26:30.384474Z digest=sha256:977af6bf658915e9b5c6807b0d2720a703d0910bb26b839198640ba4e0a26fa9

Observation ff6404e0-db9d-4ad3-9595-3a61f2f7a93e · inbound

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion cites this paper.

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-15T13:43:40.241796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:43:40.241796Z digest=sha256:b1038374ab43a113ba75c8920ff76e9cb3ca6d33da9baf88071176d39e8425c1

Observation 24d8b381-7d8d-4761-8d9d-7a3e93f2bca4 · inbound

Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision cites this paper.

Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T18:42:09.257514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:42:09.257514Z digest=sha256:c90ce1a2dd53bf8eeb5649399c5e9c7f343ff73b0577e4afef7caacd6b8b52e1

Observation aecb237b-c5fb-4829-848d-db0df7e067c6 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T08:34:56.898815Z digest=sha256:6ef99bc55de217f7eb774649db826d09a7f7da3380c84f4d480917dfe3acc87b

Observation 4edbecd6-0c16-47a1-9d78-9480c99cf225 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-21T10:44:07.726139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T10:43:27.176535Z digest=sha256:e23ee7d5b2aabffe22fe95cfcda81fcd53ca4a7ace04c1709116c0b4232bd201

Observation b84ee168-78db-49ea-a7d2-ab11bedccfc5 · inbound

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning cites this paper.

The Silent Thought: Modeling Internal Cognition in Full-Duplex Spoken Dialogue Models via Latent Reasoning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T22:57:11.059962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:57:11.059962Z digest=sha256:77d2580c2d2593b1d6d10ac0a1570c473df96e240bf070a4f3cee84bdbbe9972

Observation 7cdda8ef-ac8a-4653-9721-18588a276599 · inbound

TiCo: Time-Controllable Spoken Dialogue Model cites this paper.

TiCo: Time-Controllable Spoken Dialogue Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T00:38:52.182973Z digest=sha256:f7f0c2fd0bd308c9adf4a9ca780e13cd199c5d328c76f475af3a351075070b7a

Observation 5945c65c-a5a3-41ba-ae7d-78afb580d508 · inbound

Sharp spectral estimates for free boundary problems arising in plasma physics cites this paper.

Sharp spectral estimates for free boundary problems arising in plasma physics GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T19:56:33.793167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T19:56:33.793167Z digest=sha256:76c8c334cc6d8f3271d04b19e6d39295ed6e2ea6f1320ff4e7f376757f648a39

Observation 7e02c3ed-03eb-4df1-85c8-3945770da139 · inbound

FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection cites this paper.

FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T20:55:09.141906Z digest=sha256:cd25c69dcfd425d330c672abc110ff4c4221e227ea876daaf2fd4efe612c3c89

Observation f79f6af8-d026-4f35-a0cd-278bcafbacf1 · inbound

Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning cites this paper.

Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T10:41:51.251736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T10:41:51.251736Z digest=sha256:8c3d5f1f333e34e828a5295acff924fa0fec9d733018e902c3e61dc1df0cc3c3

Observation 52e90ac9-568a-4373-bad9-cc5b2a1fb956 · inbound

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs cites this paper.

Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T18:06:50.408402Z digest=sha256:7cd0b18714d696de423661114d150171a926dd6ea57a4488b738ce43e3b5fe01

Observation ef66c270-e501-472a-aed7-24e6d05efed4 · inbound

GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking cites this paper.

GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:40:59.993298Z digest=sha256:8ab1dd9f4de90bb5b6eb718878db19b564783942d4e6614be1b07c4480f9709e

Observation e414a280-1a0a-4bde-8880-b714bf72858e · inbound

HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models cites this paper.

HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:24:27.118694Z digest=sha256:b9fab3ce583b9e10ab3c087826b90d594543ed1c8da04bef771e9ba712f2f644

Observation d0a4d9f3-2bb2-4bd5-bb3e-8e7585d86f17 · inbound

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning cites this paper.

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T14:22:25.660785Z digest=sha256:129cfdea62c68938acd2bf8ee025eaffb8628d01553e6cf9ad3c269dddedf238

Observation 4e1f9513-42b1-4389-886f-170370dafe4b · inbound

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff cites this paper.

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T20:09:57.922788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:09:57.922788Z digest=sha256:8b80ce3d98e286996f4a3feeb15fd83525f9fbde35c0cdcf3f91bd8e2a7d70a8

Observation 3332b3a4-1a5c-4135-a0ad-e8c834f7deab · inbound

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection cites this paper.

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T11:32:10.126062Z digest=sha256:6375bcb9db1ae6a23fbf38a2067cdb9754b1e92d8d2794949066ab8ebed0f799

Observation b1ee6800-90a8-4286-8196-c135310f7c5b · inbound

Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints cites this paper.

Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T03:09:41.001660Z digest=sha256:35d99bdd4c81c2876e6f93136704ac97d3b5bc40d9b231fe84f4f30907775ff1

Observation 4b3b2aca-f622-42ec-a386-c53388929e1b · inbound

SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation cites this paper.

SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T00:39:04.303837Z digest=sha256:257956984c92c5893c611b731d24e8a297fdd0a282c5b4ec89e816fe25269202

Observation 78159a47-e9c3-4a35-b3f4-c571a44aa3bb · inbound

Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge cites this paper.

Full-Duplex Interaction in Spoken Dialogue Systems: A Comprehensive Study from the ICASSP 2026 HumDial Challenge GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T13:19:44.156822Z digest=sha256:c1c79dcadaed781eee66bd267685398dfd2418bb624a12a04d4bd8a65c29d370

Observation a3e5a554-578a-4010-801a-8c6f61e5d8d6 · inbound

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation cites this paper.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:93112afa6bcf00807e5d2402b329c7817e47553e4810a133e87acf88bc4aa9f9

Observation 088b1b3b-a760-4c2a-9f1c-c1e434bbbf05 · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T11:00:52.196039Z digest=sha256:7257b5401513c4a4f7bc05499e0bc467854b3ce4aee7b4cc3a44fcdd5005e03a

Observation ae6a8181-0131-4116-baea-9a6d8f374082 · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T00:49:26.507281Z digest=sha256:d802e9770e71f09a5676fab00e521c17018ea43f831025e8054cdef9f7ac28ce

Observation c4aa453e-1067-4b02-8b1d-d7ea0b94f867 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:42a5b359b9bfe61157f283555d4c03787b9f509dee5871315d2a732a6dea23e3

Observation b15eccef-3678-42c8-a3c9-70e638da3ff3 · inbound

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue cites this paper.

How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T03:08:55.753359Z digest=sha256:15d60b35ecd4c05be355115a61b2b1b56278dbef63837949f0691f24e47867fe

Observation 98d73a50-35cc-412a-ad0d-d42c203f2960 · inbound

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling cites this paper.

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T03:53:47.697648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-13T01:04:54.506749Z digest=sha256:0bce540363327d9f647fbe2b380953467281365def465d8fb822a875d3423987

Observation 2e492102-95ed-4ce4-a186-3f55b82241bc · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 124

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:39:49.034917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:ae184c23cf54ff8bc426bca6acc29c10055bd0b99e1c64c8a11a8f8ef57c7823

Observation 0c127363-d584-4c42-a74d-b5d003087eb3 · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-21T02:43:54.985808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T02:41:13.583493Z digest=sha256:e99dc7f89fbe80f1b92f8e9cecdd5bf84903918764fa95851c4f350e815362ab

Observation 85596773-8041-411c-8de9-8380400cefa5 · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.375297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:0c34bcb9c956caa664963545f1f71c4b33c9866dbcd750f98403619d87bec693

Observation 16a25c88-bcf3-4bab-9a14-93d06c2c3f09 · inbound

A Survey of Audio Reasoning in Multimodal Foundation Models cites this paper.

A Survey of Audio Reasoning in Multimodal Foundation Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-21T02:09:24.342869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T02:08:06.976461Z digest=sha256:3e99a49ecd1431413b6e33752fc9b8edbb6e51fde9e3be9d80e867d3354874a6

Observation edc17ada-56ba-49c9-86e3-10d9c1489f08 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:04:01.561200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:912ba9c71e81187d4629c58687f83ab610db8bf360f1a9bac1fb302fb02ba051

Observation cbba8d8f-bc45-413a-9fd7-63d48dc29d96 · inbound

Learning When to Think While Listening in Large Audio-Language Models cites this paper.

Learning When to Think While Listening in Large Audio-Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:43:50.743020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T18:37:34.409802Z digest=sha256:e2c74f941c5e71be6d0507751b6938c7e3727e9d1501e0c4efee7060999f4dca

Observation c75e5f79-e612-424b-ab5e-08c82c1d9955 · inbound

LaSR: Context-Aware Speech Recognition via Latent Reasoning cites this paper.

LaSR: Context-Aware Speech Recognition via Latent Reasoning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T19:22:34.950712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T19:12:59.843096Z digest=sha256:a495ea4462b0d3337708f38056dd79f37fb16f5e0646eed6f81052c99b86aaf0

Observation ddb2aaee-48dd-43de-8915-fc12e7d63ee2 · inbound

Sympatheia: Emotionally Adaptive Voice Assistant with Continuous Affect Conditioning cites this paper.

Sympatheia: Emotionally Adaptive Voice Assistant with Continuous Affect Conditioning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-28T18:02:27.090218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T17:56:47.512549Z digest=sha256:170c6acc8e07aa040b0b87764c3e1b2cecf3aedb8771de8fe0070e7276631696

Observation 363b0d7f-2e34-4f19-b7e8-1c3d336660ec · inbound

MOSS-Audio Technical Report cites this paper.

MOSS-Audio Technical Report GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:24.986552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:2de33924b2ec932cb4eecdca1fbf839154dbed82d8ef910c1b487faa976bd3b5

Observation 46d7b5d9-f76d-49b4-a521-2b4d25093144 · inbound

IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems cites this paper.

IRAF: Interference-Resilient Adaptive Fusion for Noise-Robust End-to-End Full-Duplex Spoken Dialogue Systems GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-02T15:37:06.809939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T23:42:38.203116Z digest=sha256:f9413943bcb19dc71a5482c561b80a78a49ca0516f880419dff91e0cad9380fd

Observation 14c18da9-d85b-446c-a9f6-ca86dbe854ac · inbound

Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models cites this paper.

Beyond Semantic Dominance: Cognitive Affective Reasoning and Empathetic Response Alignment in Audio Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:47:20.044588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T21:16:07.007759Z digest=sha256:c2cb10db61087a921466e3b2a78a18291306cd9219883f507254e7e7e227a812

Observation 51b0275c-260b-4804-a63b-b55679767853 · inbound

Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering cites this paper.

Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-03T05:07:39.514736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T13:22:52.428152Z digest=sha256:2984d728e5710b9f2d8da9808ba40e788e41053883d85f718116e86f2e9ffdca

Observation 2f7da106-8fab-442b-9c8f-1b5a96b9c611 · inbound

Benchmarking Neural Speech Compression from a Rate-Distortion Perspective cites this paper.

Benchmarking Neural Speech Compression from a Rate-Distortion Perspective GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-07-03T12:48:12.274727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T08:43:35.279033Z digest=sha256:e3d5fd533d1c7983d0d929d4e5ae65a5a9455cd67a2663faa64e14c8bbb9d0c4

Observation ed2af017-186b-431c-babe-07a59fa63cfe · inbound

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation cites this paper.

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-03T13:18:12.822482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T08:18:23.182355Z digest=sha256:ad7b0dff8400bd9d4ada17cf21aeb09631ef13d50ef59485fc3a9cb1e13ea673

Observation 2f4fe339-2ce1-4cc4-baa6-127bc09eb08f · inbound

Endpoint Anticipation for Low-Latency Spoken Dialogue cites this paper.

Endpoint Anticipation for Low-Latency Spoken Dialogue GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T16:28:39.164543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-27T05:37:17.684884Z digest=sha256:fed1caff1842c58d494c3c6881237ca5685cb20c65ba95975e99b0ca1bdadac3

Observation ba53d677-58c0-49fc-a6a8-0252b2c5af32 · inbound

Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding cites this paper.

Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:04:20.899118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-30T07:03:08.617257Z digest=sha256:83e04283bd34b222fb67f538737f99b7946bf1d7f6874f2994ab6f055972c9be

Observation dd8dc385-a372-410b-8289-a590a05290c1 · inbound

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation cites this paper.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.834436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:4207bad9fec8fc5ab0a604b0ce48d92d571c011d202371eef6660846194298cb

Observation e2d1cb7f-9588-4af1-8b32-c1d1f461a2c8 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 212

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T11:55:42.255512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:5b53c035b8d3924e7361d3df3aeef1142481470ed963784922ed57718776cc28

Observation bb37c311-8601-4b7d-a64f-07654ece0ec2 · inbound

Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis cites this paper.

Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-02T06:56:44.332417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-02T06:36:42.174254Z digest=sha256:e9fde3cee95b1eebd8195265471bb6deb1bafea928869e44b159d7544774b31d

Observation 300a1252-3d10-4733-87c3-75c4314afba4 · inbound

Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning cites this paper.

Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-03T14:38:28.596715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-03T14:35:17.004680Z digest=sha256:299a51c65125c0eacd9ec4e36e4e492ac766aca1ef85ebe6b2db68e1933382a7

Observation ee5c60a1-fc3e-4c11-939c-c96a57f55662 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 175

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.661580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:380ffb9b495ec4b778f0cfcf33c0b6974a77100ccced4d34f6e9673098980550

Observation 5b79b252-f52f-44bd-b2a5-259558dc1f95 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 175

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:b4071f0708991abe7375c24c1dcadc0471251ad3526bb7ec9c45f604c7f9c1d0

Observation c2226ca2-91bf-4957-b01f-fb74290ba53e · inbound

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs cites this paper.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.544716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:cd0e9d9c520befdce6faab8222ba9aac9e5322f85f825bc8557fe29b2fba92b7

Observation 99fb2f1a-4f93-43a9-8bd2-85cdf477a05a · inbound

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models cites this paper.

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T11:20:18.091728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:20:18.091728Z digest=sha256:53d0f438581cca3c7a65d9366f11c9ee07a005f8804797c94ffb78a5a6d45397

Observation ebb9399c-db7c-4ad2-a17d-9402744aee6b · inbound

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning cites this paper.

Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T10:40:25.933165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:40:25.933165Z digest=sha256:adb7a88ef40abcc610866078094176692e0f457369764da598ccb4514aa6aa2a

Observation 798d44f9-d625-476e-b436-eb0ef3435f70 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:15.461952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:15.461952Z digest=sha256:4c731ebcd350f53fd25075930a321dd4903e4cb39cd0585ad68b7f2a20c389af

Observation 10adce2e-e654-40fc-9766-7909b25b3abd · inbound

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation cites this paper.

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T04:20:41.694183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:20:41.694183Z digest=sha256:953cea2631521d004c1c3c9c066645e7d77ca2f658e3b570658775a605e282c7

Observation c648a9b2-30be-496d-9dd7-ddc0d543d6e7 · inbound

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects cites this paper.

DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T00:33:55.283870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:33:55.283870Z digest=sha256:5f88384f82cb2cef64a53b1f71ab38dd4421c581c5a21800c8597e86c6330313

Observation 9515d57e-b62e-4052-85c9-8c62812abc88 · inbound

EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models cites this paper.

EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T21:52:46.705383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:52:46.705383Z digest=sha256:f62af249462443713bd0784b059fb15abdd30d3ca7232955010bf87de55f442f