Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:11:43.731582Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2608.10878.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:11:43.731582Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:11:43.564832Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-12T15:11:44.295858Z
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 243e76be-3a72-4c4d-ad9e-7d712b096dfe · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction To manage these complex conversational dynamics, a responsive system must continuously estimate fine-grained turn states
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9ee589f8-f928-496a-8c0b-be97d9d7ed30 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 55cf7189-d921-4510-9de8-1e226fe42d8b · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction We further introduce an ASR-anchored supervision method that projects word-level turn annotations onto the frame-level ASR token timeline
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bc457826-56a3-4a4b-8ed3-9aa7ee5d3c48 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c6d0c0a1-80ff-4cf2-9700-f4d23e32a6b6 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction um”, “ah
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f57329ae-9658-4038-85ad-655f281767ec · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Data Preparation The corpora used in this work consist of two parts: Chinese- English ASR data and turn-taking data
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 71c5d73c-f4b7-40dc-9aa6-1d38384316e0 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 60c7699a-dfa3-44d8-aedc-da5b7668c61d · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b0516a04-94dd-4295-91eb-881aa7fb50ef · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Turn-taking in conversational systems and human- robot interaction: a review,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb58300e-019c-450b-980c-0d6af1b55519 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Generative spoken dialogue language modeling,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c42fa91-5938-4548-9fdb-cca45f2b2dd5 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Moshi: a speech-text foundation model for real-time dialogue
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccfff85c-d97c-4da1-b2b7-fa3fe2311594 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Freeze-omni: A smart and low latency speech- to-speech dialogue model with frozen LLM,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1ddfa5ec-f00b-4023-b058-9dfe69d2a306 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Omniflatten: An end-to- end gpt model for seamless voice conversation,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 73769c5b-084b-41e4-b6f0-0323e246d9c8 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Personaplex: V oice and role control for full duplex conversational speech models,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c20728d0-e494-4e2f-9e3e-4cf15a49ae8f · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c80814d9-981f-454c-8bca-679051c1e60a · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Easy turn: Integrating acoustic and lin- guistic modalities for robust turn-taking in full-duplex spoken di- alogue systems,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 42074b39-6486-459c-83bc-53b24969deee · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Jal-turn: Joint acoustic- linguistic modeling for real-time and robust turn-taking detec- tion in full-duplex spoken dialogue systems,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0cf32aa9-be7d-4cf8-af2a-3e44d287ce5f · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eca560f-6b46-4e2f-99b5-487220562900 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Soulx-duplug: Plug-and-play streaming state prediction module for realtime full-duplex speech conversa- tion,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 243f96e7-1fbd-4806-b777-5f82cac80473 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8c5eb925-300a-4967-acd2-433b8b597dd7 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Ten vad: A low-latency, lightweight and high- performance streaming voice activity detector (vad),
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2a9f1f8d-b2f3-4025-ac78-40e7dbc7ccb7 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Streaming sequence-to-sequence learning with de- layed streams modeling,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf394a1b-408c-4254-8f19-e4753b0c1cc6 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Voxtral Realtime
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2768a0fb-beb0-4816-a52e-8449723d3e61 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Qwen3 Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f16645d1-d41e-4ab4-bd14-0491269f89fa · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Aishell-1: An open- source mandarin speech corpus and a speech recognition base- line,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03782fa7-f34c-43d2-9d32-8f65f8743333 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0303442-e53b-4903-88a3-74435f2c5e76 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db800827-8d9f-4ee6-b3e0-33b6c1d05d4e · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db1c7181-a875-4722-9b7e-d790bb662aaf · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03e80c0f-56db-4b13-92b2-fa4ded69fe44 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5570088e-3a19-4197-9277-ab3ef34074e6 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Kespeech: An open source speech dataset of mandarin and its eight subdialects,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3936bbef-1e34-4374-8050-28cb2e91a93b · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Lib- rispeech: an asr corpus based on public domain audio books,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d21430d1-f07b-4d58-b5af-fc516e071a1f · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce8a115d-a805-40cb-9c65-b5096ae28c3e · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fe48ccfb-ced6-4a36-87a6-b802d8e87c82 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction V oxpopuli: A large-scale multilingual speech corpus for representation learn- ing, semi-supervised learning and interpretation,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67cefe6a-0142-42b6-9bc1-041ff6ddc039 · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction The fisher corpus: A resource for the next generations of speech-to-text
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57b29c58-0dd6-4172-b7a0-775e9bc71efb · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Qwen3-ASR Technical Report
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a837235-4f00-4319-9986-02d2c942490b · outbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Uni-asr: Unified llm-based architecture for non-streaming and streaming automatic speech recognition,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc457826-56a3-4a4b-8ed3-9aa7ee5d3c48 · inbound
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.