Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:47:14.172217Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2506.07081.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:47:14.172217Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T05:37:17.684884Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T16:28:39.127928Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 44080555-c607-49e4-a760-a70df34c9de0 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training WavChat: A Survey of Spoken Dialogue Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ca75ee5-98f1-49b8-9bce-9eed04c1a35a · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training A review of subjective scales measuring the user experience of voice assistants,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54c688bf-1a8f-4aa7-8334-91adfada3c3c · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Gemini: A Family of Highly Capable Multimodal Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f0da420-5713-49ff-9605-d978e32a8c85 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training OpenAI-gpt-4o,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 261a5a25-60ab-49de-b987-497f08b38a76 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Improved End- of-Query Detection for Streaming Speech Recognition,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation be2a5081-a9d9-4e68-9d62-720bfb6fda20 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training V oice activity projection: Self-supervised learning of turn-taking events,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 651c22f7-0f2f-4f6f-842d-169ef6f7cf64 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics ,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf9ee3b0-9255-4331-84b8-4392ca4b0ce8 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Root causes of lost time and user stress in a simple dialog system,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96915e7a-187b-4b5b-b6f9-52620342af30 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training A statistical model-based voice activity detection,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f4f2787-25e3-4410-b0d1-4d46fba5cc3d · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training A Convolutional Neural Network Smartphone App for Real-Time V oice Activity Detection,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b96e219c-d423-4ddf-9bd5-d3b65f4c9073 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Temporal modeling using dilated convolution and gating for voice-activity-detection,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7ad4af9c-0153-444c-8f67-863e23fd20bd · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Robust end-of-utterance detection for real-time speech recognition applications,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f50f5b3e-03d1-4f6b-8d10-1ec217bbbdf2 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Combining acoustic embeddings and decoding features for end-of-utterance detection in real-time far-field speech recognition systems,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5dc9aed4-5c79-4292-830b-050f5772222d · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Dynamic speech endpoint detection with regression targets,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b86fa938-e980-45c9-a34f-227472dbd7e5 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training SoundStream: An End-to-End Neural Audio Codec,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 34bd2dc7-2de0-4b9d-a23b-3ffbca6bbfe4 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training High Fidelity Neural Audio Compression,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 17e1c7d0-bc0f-4cfc-81f8-2d25bb286c4e · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Audiodec: An Open-Source Streaming High-Fidelity Neural Audio Codec,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9d41f697-b6d7-4954-95f6-dc506c121022 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Moshi: a speech-text foundation model for real-time dialogue
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a363b232-9cb9-456b-9fd4-f460d74c6255 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Codec-SUPERB: An In-Depth Analysis of Sound Codec Models,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6595c748-2f80-4c2f-b947-401edb609c36 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eae1316f-4479-4d94-9707-99e0bbf49b08 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9893ded1-943e-47cd-b054-faab046ad98c · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training High-Fidelity Simultaneous Speech-To-Speech Translation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63e01f50-c2ee-408e-a658-eb4bf3d9617c · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Self-supervised speech representation learning: A review,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a621974-0784-420e-8bea-cabd6897ef01 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training HILCodec: High-Fidelity and Lightweight Neural Audio Codec
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abbbbc0c-103e-4429-a95f-dbb4fbdbbd95 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training End-to-end speech endpoint detection utilizing acoustic and language modeling knowledge for online low- latency speech recognition,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d98b2206-b1c5-4486-9f80-464af49ccf38 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Joint endpointing and decoding with end-to-end models,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4344bebb-23e6-4ca9-83c1-9f4ab23c83a8 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Turn-Taking Prediction for Natural Conversational Speech,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1aad363c-8c0b-479f-872f-b23e539ee1a3 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Towards fast and accurate streaming end-to-end ASR,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 30dea1f6-734e-4b63-899a-fcc8d19ed347 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Streaming Automatic Speech Recognition with Re-blocking Processing Based on Integrated V oice Activity Detection
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5881084-3c11-4493-bfc4-749fdef1cbd7 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Two-Pass Endpoint Detection for Speech Recognition,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f50655af-451a-4b4e-9481-990a3529c45f · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Towards Accurate and Real-Time End-of-Speech Estimation,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 473084b1-adfe-47ea-87a7-353b44ab619d · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Text Injec- tion for Capitalization and Turn-Taking Prediction in Speech Models,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e73b4fd-9911-40ce-8881-1abafd2b4420 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Multilingual turn-taking prediction using voice activity projection,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 89e63db1-7533-4c20-be69-3ed56f734810 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Yeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 45cbefc9-f0af-4ab9-b512-200a9d8bfa12 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Turn-Taking and Backchannel Prediction with Acoustic and Large Language Model Fusion,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c4f5350-1927-4723-af50-68fd90cd8876 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training End-to-end automatic speech recognition integrated with CTC-based voice activity detection,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4debdea7-6324-49e2-aece-de6ddf6788b2 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Text- free prosody-aware generative spoken language modeling,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc82655a-cacc-48b4-b0f3-ee6fd9074f98 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Simple and controllable music generation,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43085501-452b-4b50-97a8-6d7ec417e90f · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training SpokenWOZ: a large-scale speech-text benchmark for spoken task-oriented dialogue agents,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 60401cb3-9b9e-4899-8663-ca7f1a9d073f · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Silero V AD: pre-trained enterprise-grade V oice Activity Detector (V AD), Number Detector and Language Classifier,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c425acdf-4ba8-4ec4-9a02-d5a7cb25ca4e · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training PyTorch: an imperative style, high- performance deep learning library,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc38f094-28c5-4cad-b465-4167ba175604 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Transformers: State-of-the-art natural language processing,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c25e00e3-6b18-4177-a50b-80c08dcf091c · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Modeling turn-taking in human-to-human spoken dialogue datasets us- ing self-supervised features,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2768c705-0e2b-4bd9-9b30-d63fb9d189e3 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c79f311-4e44-4233-a9bf-7396fc98ea25 · outbound
Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f8ee3e4-80fe-4a9d-bb4c-5c66b4c1ee93 · inbound
Endpoint Anticipation for Low-Latency Spoken Dialogue Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.