Pith. sign in

Paper Citation Record · LEDGER

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction

As of 12 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2608.10878.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10878 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:11:43.731582Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:11:43.564832Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T15:11:44.295858Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact3
  • verified fuzzy12
  • unresolved22
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 243e76be-3a72-4c4d-ad9e-7d712b096dfe · outbound

This paper cites To manage these complex conversational dynamics, a responsive system must continuously estimate fine-grained turn states.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction To manage these complex conversational dynamics, a responsive system must continuously estimate fine-grained turn states

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.614557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.549766Z digest=sha256:4b150fea1eccd90ff46cf577651eefdcd5937a94710db0a2308c00b9795d4e74

Observation 9ee589f8-f928-496a-8c0b-be97d9d7ed30 · outbound

This paper cites an unresolved cited work.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:11:44.597022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.555407Z digest=sha256:9a6511529263d68c96a38911622d747202b0a02562cfd8b3db080843581cc30f

Observation 55cf7189-d921-4510-9de8-1e226fe42d8b · outbound

This paper cites We further introduce an ASR-anchored supervision method that projects word-level turn annotations onto the frame-level ASR token timeline.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction We further introduce an ASR-anchored supervision method that projects word-level turn annotations onto the frame-level ASR token timeline

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.578653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.559896Z digest=sha256:ef4833a2cded1e34d4ac0e4332e8856c27803d57cb6210f5fcbb304fd7834f33

Observation bc457826-56a3-4a4b-8ed3-9aa7ee5d3c48 · outbound

This paper cites X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:11:44.300966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.564832Z digest=sha256:56d0858b1bf583037f90dc9b7f163c51d9184b074c7c095029a4033a3dedc50a

Observation c6d0c0a1-80ff-4cf2-9700-f4d23e32a6b6 · outbound

This paper cites um”, “ah.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction um”, “ah

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.561641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.570788Z digest=sha256:85049e9665e0128d25e2ac722441c417324b67a4c0ac0986d478a316413c00ba

Observation f57329ae-9658-4038-85ad-655f281767ec · outbound

This paper cites Data Preparation The corpora used in this work consist of two parts: Chinese- English ASR data and turn-taking data.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Data Preparation The corpora used in this work consist of two parts: Chinese- English ASR data and turn-taking data

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.546774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.576540Z digest=sha256:9ac3dd2773b2ab3e923ee083eeac63f8c8b776ae47ef1867b930b94d338230f5

Observation 71c5d73c-f4b7-40dc-9aa6-1d38384316e0 · outbound

This paper cites an unresolved cited work.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Unresolved cited work

Reference 7

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T15:11:44.531109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.581668Z digest=sha256:6f7dea6eb33e9fe580dec331ba125c8b37b2e83ef26eb0874795600f3657b51b

Observation 60c7699a-dfa3-44d8-aedc-da5b7668c61d · outbound

This paper cites an unresolved cited work.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:11:44.515454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.586925Z digest=sha256:d20d63950f7fba2bfb14461864b5ea225ad2d5931e4b96325da6e26b7a7c4456

Observation b0516a04-94dd-4295-91eb-881aa7fb50ef · outbound

This paper cites Turn-taking in conversational systems and human- robot interaction: a review,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Turn-taking in conversational systems and human- robot interaction: a review,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.591372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.591372Z digest=sha256:4f49c09718481a43cbe2427bdb72bd7e879a3453aec74740e7f01db0f4357725

Observation fb58300e-019c-450b-980c-0d6af1b55519 · outbound

This paper cites Generative spoken dialogue language modeling,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Generative spoken dialogue language modeling,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.595818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.595818Z digest=sha256:cafad81e35e52e1abe5a1a06a8d13cfeb7fbe4f6851792ce4de65f5831302074

Observation 8c42fa91-5938-4548-9fdb-cca45f2b2dd5 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Moshi: a speech-text foundation model for real-time dialogue

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.600630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.600630Z digest=sha256:e58a6ccddef280ab051c8408cdd98e0b0f3169fcfd566692b9ad802c5672d8e5

Observation ccfff85c-d97c-4da1-b2b7-fa3fe2311594 · outbound

This paper cites Freeze-omni: A smart and low latency speech- to-speech dialogue model with frozen LLM,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Freeze-omni: A smart and low latency speech- to-speech dialogue model with frozen LLM,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.479011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.605914Z digest=sha256:af28273fa74c94a7dc2979f7ee75f17155d246e0c290370a2b36d91fb10e54fc

Observation 1ddfa5ec-f00b-4023-b058-9dfe69d2a306 · outbound

This paper cites Omniflatten: An end-to- end gpt model for seamless voice conversation,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Omniflatten: An end-to- end gpt model for seamless voice conversation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.462822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.610969Z digest=sha256:3fb0b29734e10d93c10f2dff9200486cce09d9b01af1d0acabba1d31ce7e4702

Observation 73769c5b-084b-41e4-b6f0-0323e246d9c8 · outbound

This paper cites Personaplex: V oice and role control for full duplex conversational speech models,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Personaplex: V oice and role control for full duplex conversational speech models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.447867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.616181Z digest=sha256:27fc2047fe7438722226dacda73e25ff429ffe0f44caac8a7a8737eac0b1bb3a

Observation c20728d0-e494-4e2f-9e3e-4cf15a49ae8f · outbound

This paper cites FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.620558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.620558Z digest=sha256:6add2981b4cb97828cbd3378c41f53fca2e249b92d2ec7a9f3a48fc3b02bb1d2

Observation c80814d9-981f-454c-8bca-679051c1e60a · outbound

This paper cites Easy turn: Integrating acoustic and lin- guistic modalities for robust turn-taking in full-duplex spoken di- alogue systems,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Easy turn: Integrating acoustic and lin- guistic modalities for robust turn-taking in full-duplex spoken di- alogue systems,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.432904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.625434Z digest=sha256:f45b5364777c1725d9879cc5e73605a03c87315dc6338182c821c7172697907f

Observation 42074b39-6486-459c-83bc-53b24969deee · outbound

This paper cites Jal-turn: Joint acoustic- linguistic modeling for real-time and robust turn-taking detec- tion in full-duplex spoken dialogue systems,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Jal-turn: Joint acoustic- linguistic modeling for real-time and robust turn-taking detec- tion in full-duplex spoken dialogue systems,

Reference 17

Resolution
verified exact
raw_fallback, observed 2026-08-12T15:11:44.247796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.629922Z digest=sha256:9002d86831ab3b2c852954d7dc4efc708aecae7dbc340d0ef611102ddf19e8ce

Observation 0cf32aa9-be7d-4cf8-af2a-3e44d287ce5f · outbound

This paper cites FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.634685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.634685Z digest=sha256:7f927f810cf5b0a1e51b5549680367def8e613a4908cd84d5ca5081a44c50034

Observation 3eca560f-6b46-4e2f-99b5-487220562900 · outbound

This paper cites Soulx-duplug: Plug-and-play streaming state prediction module for realtime full-duplex speech conversa- tion,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Soulx-duplug: Plug-and-play streaming state prediction module for realtime full-duplex speech conversa- tion,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.639915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.639915Z digest=sha256:52e5db99061918372fcf2550b954b7daa2c0049621c6a6cbc0971f0690ed0a4c

Observation 243f96e7-1fbd-4806-b777-5f82cac80473 · outbound

This paper cites JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:11:44.080022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.644343Z digest=sha256:ffd713e57e8a18505cd9051b8e281529ae4eca7f2871d196a9d6c6c53fecc904

Observation 8c5eb925-300a-4967-acd2-433b8b597dd7 · outbound

This paper cites Ten vad: A low-latency, lightweight and high- performance streaming voice activity detector (vad),.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Ten vad: A low-latency, lightweight and high- performance streaming voice activity detector (vad),

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.416760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.648947Z digest=sha256:1fb2785a37ed55f0c5f07846db7fccbd8b68a8f50e3369bf57ac08938d75b1d6

Observation 2a9f1f8d-b2f3-4025-ac78-40e7dbc7ccb7 · outbound

This paper cites Streaming sequence-to-sequence learning with de- layed streams modeling,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Streaming sequence-to-sequence learning with de- layed streams modeling,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.654024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.654024Z digest=sha256:4dc25a8887e4e4b3e8d5c9647146bbcf570d826a6f5e647defaa91c606070b22

Observation cf394a1b-408c-4254-8f19-e4753b0c1cc6 · outbound

This paper cites Voxtral Realtime.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Voxtral Realtime

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.658756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.658756Z digest=sha256:e81df3a1b8eff8df25362abf950469f8e2d09a4b06c5d02aba46aa3c03787f2b

Observation 2768a0fb-beb0-4816-a52e-8449723d3e61 · outbound

This paper cites Qwen3 Technical Report.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Qwen3 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.664098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.664098Z digest=sha256:a7fd1671561c1bbaef091c8abd96ff358084dfd0396d0f818eb5886a823f616c

Observation f16645d1-d41e-4ab4-bd14-0491269f89fa · outbound

This paper cites Aishell-1: An open- source mandarin speech corpus and a speech recognition base- line,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Aishell-1: An open- source mandarin speech corpus and a speech recognition base- line,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.669914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.669914Z digest=sha256:d6a5be63f7ee73685aa34c6ff32b08386a697734e10723d3547bc2cd963bb273

Observation 03782fa7-f34c-43d2-9d32-8f65f8743333 · outbound

This paper cites AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.674457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.674457Z digest=sha256:543cecbe31f7417745344d4425675bc8403f394161ef888b548a0e742c2829e0

Observation d0303442-e53b-4903-88a3-74435f2c5e76 · outbound

This paper cites AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.679417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.679417Z digest=sha256:a5669c438efcfe1a17ae2df96b6b362b17123903b5917902fae1d9a703589f11

Observation db800827-8d9f-4ee6-b3e0-33b6c1d05d4e · outbound

This paper cites AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.684284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.684284Z digest=sha256:5c5da2e4d1a8930f50bb6df45678f8066e9af83e865e98c79418ff2afc203b58

Observation db1c7181-a875-4722-9b7e-d790bb662aaf · outbound

This paper cites M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.689721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.689721Z digest=sha256:a244a9aed8eeba37cbc9d50ac86a98ce18c1b170158f25d56f2aa57004bd3398

Observation 03e80c0f-56db-4b13-92b2-fa4ded69fe44 · outbound

This paper cites Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.379233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.694583Z digest=sha256:37b8c874268c89f3645126786ac21f7a029c3cad10f0c615713fc080fa74981c

Observation 5570088e-3a19-4197-9277-ab3ef34074e6 · outbound

This paper cites Kespeech: An open source speech dataset of mandarin and its eight subdialects,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Kespeech: An open source speech dataset of mandarin and its eight subdialects,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.362696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.698968Z digest=sha256:0354f09579d20bd4686fc1bbf8cda83daced44e5ae01bc568640f594aae0989d

Observation 3936bbef-1e34-4374-8050-28cb2e91a93b · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Lib- rispeech: an asr corpus based on public domain audio books,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.703973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.703973Z digest=sha256:8f9b34a5b4642282f9a2477a710c6666e475977702c87cf5c0c89406327b84d3

Observation d21430d1-f07b-4d58-b5af-fc516e071a1f · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.708566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.708566Z digest=sha256:b20026cc12fa7fec1da6e32f0cfddd992e630bea20ec24545b1bb1c3933886e4

Observation ce8a115d-a805-40cb-9c65-b5096ae28c3e · outbound

This paper cites Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.336181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.713416Z digest=sha256:c851eb07effb214b97b84dd69aabcada30c6ccc8e37958e1fa378e2a97cc5fec

Observation fe48ccfb-ced6-4a36-87a6-b802d8e87c82 · outbound

This paper cites V oxpopuli: A large-scale multilingual speech corpus for representation learn- ing, semi-supervised learning and interpretation,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction V oxpopuli: A large-scale multilingual speech corpus for representation learn- ing, semi-supervised learning and interpretation,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.717739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.717739Z digest=sha256:58c8bf583e58a99b1c4d94697399bfc5053c0c63b8ba7bc2231a236bc6adea9f

Observation 67cefe6a-0142-42b6-9bc1-041ff6ddc039 · outbound

This paper cites The fisher corpus: A resource for the next generations of speech-to-text.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction The fisher corpus: A resource for the next generations of speech-to-text

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.722161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.722161Z digest=sha256:72459defa67a63c1eee63471dad3ad911bf0b3eb56ad3093a2115c6c1e3a2844

Observation 57b29c58-0dd6-4172-b7a0-775e9bc71efb · outbound

This paper cites Qwen3-ASR Technical Report.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Qwen3-ASR Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.726618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.726618Z digest=sha256:f48ea41642ba0a4e9d4a1d9ed7dde190a383f11b09e321e65725949065aed4f7

Observation 3a837235-4f00-4319-9986-02d2c942490b · outbound

This paper cites Uni-asr: Unified llm-based architecture for non-streaming and streaming automatic speech recognition,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Uni-asr: Unified llm-based architecture for non-streaming and streaming automatic speech recognition,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.731582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.731582Z digest=sha256:9030eedecefa47ded310d04029a9d18da89bd80389056dc37494db72802adc51

Pith citing papers

Observation bc457826-56a3-4a4b-8ed3-9aa7ee5d3c48 · inbound

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction cites this paper.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:11:44.300966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.564832Z digest=sha256:56d0858b1bf583037f90dc9b7f163c51d9184b074c7c095029a4033a3dedc50a