Pith. sign in

Paper Citation Record · LEDGER

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction

As of 12 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2608.10878.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10878 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:11:43.731582Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:11:43.564832Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T15:11:44.295858Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact3
  • verified fuzzy12
  • unresolved22
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 243e76be-3a72-4c4d-ad9e-7d712b096dfe · outbound

This paper cites To manage these complex conversational dynamics, a responsive system must continuously estimate fine-grained turn states.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction To manage these complex conversational dynamics, a responsive system must continuously estimate fine-grained turn states

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.614557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.549766Z digest=sha256:31b7835ef7a97a51adc8023fc562aff555b7e71f8b70d5d9e41763dc64d0ab89

Observation 9ee589f8-f928-496a-8c0b-be97d9d7ed30 · outbound

This paper cites an unresolved cited work.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:11:44.597022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.555407Z digest=sha256:431112eabce08a9fe3eae77ee7e903a3047593f875ae95775e2cc1d155d81b14

Observation 55cf7189-d921-4510-9de8-1e226fe42d8b · outbound

This paper cites We further introduce an ASR-anchored supervision method that projects word-level turn annotations onto the frame-level ASR token timeline.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction We further introduce an ASR-anchored supervision method that projects word-level turn annotations onto the frame-level ASR token timeline

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.578653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.559896Z digest=sha256:752b207e2bc3b0031060a9bd81bfe1813c6f7b549715c622b2ee52855abc843f

Observation bc457826-56a3-4a4b-8ed3-9aa7ee5d3c48 · outbound

This paper cites X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:11:44.300966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.564832Z digest=sha256:ae15fe8ad5eaff5aaf08afca52f8e7e962e207f23c3d57272d219c61bf343af0

Observation c6d0c0a1-80ff-4cf2-9700-f4d23e32a6b6 · outbound

This paper cites um”, “ah.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction um”, “ah

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.561641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.570788Z digest=sha256:bdbbc8a1b955334b24e136f5805e41b78054151a032c15cc45133c26f44f8e17

Observation f57329ae-9658-4038-85ad-655f281767ec · outbound

This paper cites Data Preparation The corpora used in this work consist of two parts: Chinese- English ASR data and turn-taking data.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Data Preparation The corpora used in this work consist of two parts: Chinese- English ASR data and turn-taking data

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.546774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.576540Z digest=sha256:4833bd79afbb836f54364241054b041390b531d5d3ad959d5b2d8a2ae06aabb1

Observation 71c5d73c-f4b7-40dc-9aa6-1d38384316e0 · outbound

This paper cites an unresolved cited work.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Unresolved cited work

Reference 7

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T15:11:44.531109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.581668Z digest=sha256:576047e12c18abdddd1c64512ede57e4dc6dbd1a2492cb29f4eb8a3562c65752

Observation 60c7699a-dfa3-44d8-aedc-da5b7668c61d · outbound

This paper cites an unresolved cited work.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:11:44.515454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.586925Z digest=sha256:c756f5076c09dfa9ecd476dd413417caf7f082804f85043425ee7603bcfdcea8

Observation b0516a04-94dd-4295-91eb-881aa7fb50ef · outbound

This paper cites Turn-taking in conversational systems and human- robot interaction: a review,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Turn-taking in conversational systems and human- robot interaction: a review,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.591372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.591372Z digest=sha256:49104dbf82acacbdf9eb9d002d29b3a1241f08835de5e826ef2410e0dd5d5c0c

Observation fb58300e-019c-450b-980c-0d6af1b55519 · outbound

This paper cites Generative spoken dialogue language modeling,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Generative spoken dialogue language modeling,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.595818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.595818Z digest=sha256:d506cfd1cac7a85c1549bf1a539b0b073e1cc5d77b5b235627762e14158213d1

Observation 8c42fa91-5938-4548-9fdb-cca45f2b2dd5 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Moshi: a speech-text foundation model for real-time dialogue

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.600630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.600630Z digest=sha256:7f4dfcfcd2cfdd7cdf435b50f130f745e6898ce2ddff20207f17c17f2f841995

Observation ccfff85c-d97c-4da1-b2b7-fa3fe2311594 · outbound

This paper cites Freeze-omni: A smart and low latency speech- to-speech dialogue model with frozen LLM,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Freeze-omni: A smart and low latency speech- to-speech dialogue model with frozen LLM,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.479011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.605914Z digest=sha256:91dc0a4370c1f5c6d2129f93c6059333d541d633b7263ac4990610d08d5cf6e7

Observation 1ddfa5ec-f00b-4023-b058-9dfe69d2a306 · outbound

This paper cites Omniflatten: An end-to- end gpt model for seamless voice conversation,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Omniflatten: An end-to- end gpt model for seamless voice conversation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.462822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.610969Z digest=sha256:72262cf0c91d4e25abafb7a7af51ab6318453c6cd1ddefeea7ceb2525ea2730b

Observation 73769c5b-084b-41e4-b6f0-0323e246d9c8 · outbound

This paper cites Personaplex: V oice and role control for full duplex conversational speech models,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Personaplex: V oice and role control for full duplex conversational speech models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.447867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.616181Z digest=sha256:129318a98d93a08a351ba90c3d5f84066b6b8d3616a0dfe946c9ca049dbd1137

Observation c20728d0-e494-4e2f-9e3e-4cf15a49ae8f · outbound

This paper cites FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.620558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.620558Z digest=sha256:be432bdcdf9d42e6dd635d730cf9a7848567e0fd6577d61bc582cf3f83049f5c

Observation c80814d9-981f-454c-8bca-679051c1e60a · outbound

This paper cites Easy turn: Integrating acoustic and lin- guistic modalities for robust turn-taking in full-duplex spoken di- alogue systems,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Easy turn: Integrating acoustic and lin- guistic modalities for robust turn-taking in full-duplex spoken di- alogue systems,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.432904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.625434Z digest=sha256:3afaae455c39dd732d06aa9eeceefb619b216acdce09bfc9f8383a7187ec95c6

Observation 42074b39-6486-459c-83bc-53b24969deee · outbound

This paper cites Jal-turn: Joint acoustic- linguistic modeling for real-time and robust turn-taking detec- tion in full-duplex spoken dialogue systems,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Jal-turn: Joint acoustic- linguistic modeling for real-time and robust turn-taking detec- tion in full-duplex spoken dialogue systems,

Reference 17

Resolution
verified exact
raw_fallback, observed 2026-08-12T15:11:44.247796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.629922Z digest=sha256:2b12bf0b8730dc4c317bef8148b24c151364a06d50bbcc4d4581366048b6354c

Observation 0cf32aa9-be7d-4cf8-af2a-3e44d287ce5f · outbound

This paper cites FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.634685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.634685Z digest=sha256:cc3108357df5590471f7dfd79d92459a670473fa1dc1eb2206625fc5c8cca734

Observation 3eca560f-6b46-4e2f-99b5-487220562900 · outbound

This paper cites Soulx-duplug: Plug-and-play streaming state prediction module for realtime full-duplex speech conversa- tion,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Soulx-duplug: Plug-and-play streaming state prediction module for realtime full-duplex speech conversa- tion,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.639915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.639915Z digest=sha256:6a0ef8448c590912d21d08d5f65765ba8c1126f13b65a32f182ea6eb5026735d

Observation 243f96e7-1fbd-4806-b777-5f82cac80473 · outbound

This paper cites JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:11:44.080022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.644343Z digest=sha256:14163d0fb958e3037969d488237dcdbd671bbe4ab464cc1a9886e73d77558008

Observation 8c5eb925-300a-4967-acd2-433b8b597dd7 · outbound

This paper cites Ten vad: A low-latency, lightweight and high- performance streaming voice activity detector (vad),.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Ten vad: A low-latency, lightweight and high- performance streaming voice activity detector (vad),

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.416760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.648947Z digest=sha256:dfd7d86f1cdb69ed050a692c77fa34f9009f3d42f076fe2c4515b73d563dbdf4

Observation 2a9f1f8d-b2f3-4025-ac78-40e7dbc7ccb7 · outbound

This paper cites Streaming sequence-to-sequence learning with de- layed streams modeling,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Streaming sequence-to-sequence learning with de- layed streams modeling,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.654024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.654024Z digest=sha256:af457ff4b916fb6113394c2aa640816c109adced80cba610cb1ce45044a40cc2

Observation cf394a1b-408c-4254-8f19-e4753b0c1cc6 · outbound

This paper cites Voxtral Realtime.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Voxtral Realtime

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.658756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.658756Z digest=sha256:d2001273deb7d3089f5e05db92a2e1ec22975c130c915082588603cf4fd1b9e9

Observation 2768a0fb-beb0-4816-a52e-8449723d3e61 · outbound

This paper cites Qwen3 Technical Report.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Qwen3 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.664098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.664098Z digest=sha256:8fe177cbeee87d5e32164947dedf9f16b3c0fe39ceae6ec1f2786ed59bfe53de

Observation f16645d1-d41e-4ab4-bd14-0491269f89fa · outbound

This paper cites Aishell-1: An open- source mandarin speech corpus and a speech recognition base- line,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Aishell-1: An open- source mandarin speech corpus and a speech recognition base- line,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.669914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.669914Z digest=sha256:f49c166cd45e45059fd72ff5aeef6170c2cb7a4e26d74c2e796cd0aca35ed801

Observation 03782fa7-f34c-43d2-9d32-8f65f8743333 · outbound

This paper cites AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.674457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.674457Z digest=sha256:5b44117b228aad0efa4e5a1b8446eb58d30d3d9f22f32dd003c222629cc41a70

Observation d0303442-e53b-4903-88a3-74435f2c5e76 · outbound

This paper cites AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.679417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.679417Z digest=sha256:3651a214bd3bffddfe32bf1bbcdbd640ec5d77b70f2785fdae53fc9257e53856

Observation db800827-8d9f-4ee6-b3e0-33b6c1d05d4e · outbound

This paper cites AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction AISHELL-4: An Open Source Dataset for Speech Enhancement, Separation, Recognition and Speaker Diarization in Conference Scenario

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.684284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.684284Z digest=sha256:d61ed3429432e116764bd98db1e043176f5730c1b29c6072f627b91c0638fb5a

Observation db1c7181-a875-4722-9b7e-d790bb662aaf · outbound

This paper cites M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction M2met: The icassp 2022 multi- channel multi-party meeting transcription challenge,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.689721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.689721Z digest=sha256:bd50febf340accc3d440cd39bd320fd31dbe4e54202af8e37b4a1f4e55d6ee9b

Observation 03e80c0f-56db-4b13-92b2-fa4ded69fe44 · outbound

This paper cites Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Wenetspeech: A 10000+ hours multi- domain mandarin corpus for speech recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.379233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.694583Z digest=sha256:362e7630c3915dccdf4a1eab14041087f3f6faa0784853ca9b504b24533440f9

Observation 5570088e-3a19-4197-9277-ab3ef34074e6 · outbound

This paper cites Kespeech: An open source speech dataset of mandarin and its eight subdialects,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Kespeech: An open source speech dataset of mandarin and its eight subdialects,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.362696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.698968Z digest=sha256:a7e367545fddd7c870df8c91860b7a1a6d1061a10816e008c72b48cee1c7991e

Observation 3936bbef-1e34-4374-8050-28cb2e91a93b · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Lib- rispeech: an asr corpus based on public domain audio books,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.703973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.703973Z digest=sha256:3d18a8861b2a957ca03d90cb12ec7cc723e6eebc8165b433ef065afc7ef0f8be

Observation d21430d1-f07b-4d58-b5af-fc516e071a1f · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.708566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.708566Z digest=sha256:aa915626a01cab49b2390c8c847df1a5424c4ca4774756f82583bb1d6662bab3

Observation ce8a115d-a805-40cb-9c65-b5096ae28c3e · outbound

This paper cites Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:11:44.336181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.713416Z digest=sha256:486a30b404a2149932ac6eb3acd539974ca1a5e0965e965f9ffaba913edc8dc6

Observation fe48ccfb-ced6-4a36-87a6-b802d8e87c82 · outbound

This paper cites V oxpopuli: A large-scale multilingual speech corpus for representation learn- ing, semi-supervised learning and interpretation,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction V oxpopuli: A large-scale multilingual speech corpus for representation learn- ing, semi-supervised learning and interpretation,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.717739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.717739Z digest=sha256:97623bdf549fd8f13437becbef26d5da27c36999c57185668df60e8a829537d3

Observation 67cefe6a-0142-42b6-9bc1-041ff6ddc039 · outbound

This paper cites The fisher corpus: A resource for the next generations of speech-to-text.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction The fisher corpus: A resource for the next generations of speech-to-text

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.722161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.722161Z digest=sha256:a4d2924150d118032ae0b6701062844f3ca85d33212cd4c53935f19183a5e8af

Observation 57b29c58-0dd6-4172-b7a0-775e9bc71efb · outbound

This paper cites Qwen3-ASR Technical Report.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Qwen3-ASR Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.726618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.726618Z digest=sha256:50ece5a32dc52a39038b0ee2062ce8685bc222e021778bc6ba5213276b699d5e

Observation 3a837235-4f00-4319-9986-02d2c942490b · outbound

This paper cites Uni-asr: Unified llm-based architecture for non-streaming and streaming automatic speech recognition,.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction Uni-asr: Unified llm-based architecture for non-streaming and streaming automatic speech recognition,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:11:43.731582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:11:43.731582Z digest=sha256:c79462661de56cebe6bcf8cedf016fd14662267bf547448f8c729ae25121c469

Pith citing papers

Observation bc457826-56a3-4a4b-8ed3-9aa7ee5d3c48 · inbound

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction cites this paper.

X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:11:44.300966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:11:43.564832Z digest=sha256:ae15fe8ad5eaff5aaf08afca52f8e7e962e207f23c3d57272d219c61bf343af0