Pith. sign in

Paper Citation Record · LEDGER

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action

As of 11 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 5 inbound Pith citation observations for arXiv:2605.20755.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.20755 v2

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T17:32:58.848455Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T10:13:57.633052Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T16:49:57.633553Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact34
  • verified fuzzy3
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9835947-2422-4c8e-b1ac-35d4d080d17c · outbound

This paper cites A Full-duplex Speech Dialogue Scheme Based On Large Language Models.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action A Full-duplex Speech Dialogue Scheme Based On Large Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.341245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:74e374f1d00955058ee3fe750e4373a561cbb71ab28be8d2b4f8b352de253747

Observation 15b4544e-2b43-4746-92dc-048b86451097 · outbound

This paper cites Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:34:57.338471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:43c5534cc6d8b0b4ff1792b2071bd91699b767971e8461835c945185743f6884

Observation 023cbe8c-e963-4b83-9c82-cdeddd7a5d96 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Moshi: a speech-text foundation model for real-time dialogue

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.332842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:078546a8875afd03509e0f2f1938d1bd0147df104f7d84cb97a5d226ec981ff2

Observation 9ba951a5-3065-44ce-8f9f-83ddbaf3159d · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.380968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:fd608a8993732a6e74bd6013a68902fdf62674d7119d0239aa3fae6915d3bb34

Observation 1aeb9e24-244b-4c88-8f5b-f3488cb8ca2c · outbound

This paper cites SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.330308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:b1b898e74363da4da91e0b7baaa1779ad31a832ab0c9b8db6e949c7727cad2bb

Observation 5faf939b-3d24-411d-a178-06f698bc79e6 · outbound

This paper cites Efficient and direct duplex 14 DuplexSLA modeling for speech-to-speech language model.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Efficient and direct duplex 14 DuplexSLA modeling for speech-to-speech language model

Reference 6

Resolution
verified exact
doi, observed 2026-06-30T17:34:56.831466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:2cd26e7e093d73056506c225f41e1f7c706ca362c9f641ad1242fa5e65ce03c0

Observation d56001ab-b8aa-44cc-8d51-70a2ce30bc94 · outbound

This paper cites Covo-audio technical report.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Covo-audio technical report

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.352574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:062fa3f6a3a1f34d19709e174249d8ebbfd6302f1db37093c693167f381fd2dd

Observation 545c6986-d639-4474-a94d-3155d96c129d · outbound

This paper cites Personaplex: V oice and role control for full duplex conversational speech models.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Personaplex: V oice and role control for full duplex conversational speech models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.358844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:60b64a2e2fe95e12e70c213b16e7c05ef54b9d4e6f360102bf0de3413ad3b842

Observation 55316a08-8ff2-4fd0-8a67-f2941ebc47ac · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.357997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:4e1f93f9cc9015270611fe7ca59092b923ae35d0d50b701267d1f8a1a1fe465f

Observation 22db588d-7ad2-4e26-87bb-d8e950105eda · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.307538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:ee28f70ecd91fee1b0ff7dfd1fb0c4e6375d0580490ac0b7fa96497cb99ace7a

Observation 4838ffaa-416f-4089-8ce4-244f17c814de · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.361577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:9469c66e997f44b82d64fba9e1d86c35118de7aa1a85fe87ca8cc0b019073f7e

Observation cf4b8bca-23fe-4d39-bbd1-7b9578862be6 · outbound

This paper cites OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:34:57.341812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:fb3d8fd1de62e487ff7c80dac689c5fc951809bfd3f5a2545c8bd03a7016143d

Observation 4c621460-4a03-419b-8118-66c9c62afaf3 · outbound

This paper cites Spirit LM: Interleaved Spoken and Written Language Model.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Spirit LM: Interleaved Spoken and Written Language Model

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.346754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:b77d17847978c4bb13f6bcd5e4b87c641941b6ce3bcd8d039009ab7141ab3d01

Observation 57dc6c13-b400-45cb-822b-464e6734da8d · outbound

This paper cites Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.349819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:5feee794b2d231340b911ecd0af4f6bfe25dbb8bf89ba2db3ac00435fe6a029e

Observation 3aee9618-5a94-4ebe-b7de-959b4f1fa015 · outbound

This paper cites Chronological thinking in full-duplex spoken dialogue language models.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Chronological thinking in full-duplex spoken dialogue language models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.344308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:ac75c8881f85d20975787ab9ef67f66c8ae5e50da262d3b0dfab0bfcecf00399

Observation 8df56ad0-10a2-47d9-b38b-ac13e967fda8 · outbound

This paper cites Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.336753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:8ee6c64c9246ca2b13c8f218a7c5c92d0cb3c94d4442545a1b425c2da8a72ee6

Observation 05fc70c6-7f92-4365-b72b-255ff4dffc37 · outbound

This paper cites Qwen2 Technical Report.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Qwen2 Technical Report

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.329158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:3c105ecdc7e788659f74cfdac5d945447cc9ec68753fa479f9b58e313998fd42

Observation 37c4dd72-d28a-40ad-9a4f-d3fb258f652b · outbound

This paper cites Qwen2-Audio Technical Report.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Qwen2-Audio Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.324473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:55516c1b1d40150e60eee0d87ac8b4f24a06e5607efbd799550ebc2d77789d90

Observation b8d0dbc3-3147-49a9-84a8-58631c65090c · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.326987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:742366fc65efaac0e0026cd4b77546948835ae235cdf5428cd7a017c7e5d7cba

Observation 0ca85e37-9c8e-4e52-9c37-61d74f64ebdd · outbound

This paper cites Step-Audio 2 Technical Report.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Step-Audio 2 Technical Report

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.386787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:fadef4ea870d0f099a4ffaa6ab2f5181340a7d1b2811c2058b42dd4a41aec60d

Observation 5c040669-e168-40ea-a788-d5918f790e6d · outbound

This paper cites Step-audio-r1 technical report.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Step-audio-r1 technical report

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.352865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:39956ae5f3bfaa12f6ee6c141ad37b9305fe8464c462e5bdc2d31b52156ba825

Observation c3fd87af-b37b-4061-a6db-076145f801fe · outbound

This paper cites Step-Audio-R1.5 Technical Report.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Step-Audio-R1.5 Technical Report

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.321667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:a0f62d25a066c84e51744448333333773da95356ad21cab25e63d06f262fef07

Observation 74e72d1d-808c-4c53-b168-6bf51ef6c5f6 · outbound

This paper cites GPT-4o System Card.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action GPT-4o System Card

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.309897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:1cc15f15b29e72336f2b48b401650819ca19a3c4e894fc9581a8cee3d8f95ff7

Observation 461db628-4cb9-4179-b028-40472f2286c7 · outbound

This paper cites Qwen3-Omni Technical Report.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Qwen3-Omni Technical Report

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.292528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:89485b22e6fb82a2bdc21e7bfc877a748730c0aeb6ea8719429b67bf41383abf

Observation 85596773-8041-411c-8de9-8380400cefa5 · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.375297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:230ef41c707527fb56f7f5e3b98c61357c9a7867356eb8502e187cd0a29016f0

Observation 708f9c9f-fc47-4a3a-b3c8-8d5fd6be44d7 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.369814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:4541e0d247bfa0d3220df310b3b1f633417b56c8fa34938e63f9892984aaa23e

Observation c011e9e1-acae-470e-9b73-c73577fff584 · outbound

This paper cites Vita-audio: Fast interleaved cross-modal to- ken generation for efficient large speech-language model.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Vita-audio: Fast interleaved cross-modal to- ken generation for efficient large speech-language model

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.298264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:1706675e9557df182a168be9a3960faa1a87a2f1b45d401efa848ecc5d4be2c0

Observation 4cea8c54-3247-4bd1-8e25-a6dcb9565d86 · outbound

This paper cites Minicpm-o 2.6: A gemini 2.5 flash level mllm for vision, speech, and full-duplex multimodal live streaming on your phone.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Minicpm-o 2.6: A gemini 2.5 flash level mllm for vision, speech, and full-duplex multimodal live streaming on your phone

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T08:54:49.620945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:2243ebea0a48c8eab68885e52bd71af94c99ad3f2e2cd8cdbf1aed02c30893ba

Observation 00731764-30ef-4d71-9af2-8a743e17e4b5 · outbound

This paper cites Mamba in speech: Towards an alternative to self-attention.IEEE Transactions on Audio, Speech and Language Processing.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Mamba in speech: Towards an alternative to self-attention.IEEE Transactions on Audio, Speech and Language Processing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T08:54:49.618089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:a9c50751ff34f373eb161be1d3e43b0e938973ed3b65cd232dac7fa25d8859d8

Observation c714de8a-37f2-4928-b7fb-a53f225d7270 · outbound

This paper cites Code-switching speech recognition under the lens: Model-and data-centric perspectives.IEEE Transactions on Audio, Speech and Language Processing.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Code-switching speech recognition under the lens: Model-and data-centric perspectives.IEEE Transactions on Audio, Speech and Language Processing

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T08:54:49.623697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:150bfe01578f4cfe7952b415f62a8917f063e8009a463f7865d3a36900097f86

Observation 6206bdeb-4f09-47b2-8699-d386e2731871 · outbound

This paper cites Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.378101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:97a5046cd9269a1e74fb238d209294a81795ea73ba46bc8038b61e58cfa1e51e

Observation 34394e45-5b0c-431d-baca-212bcd26ae95 · outbound

This paper cites Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:34:57.367281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:8801b241d5889e7608a9ae15c5f626fa3f9d6941f7aadc6ec2944306ce356830

Observation cdf6a87b-8d9e-45d1-9503-e0403f078583 · outbound

This paper cites Wildspeech-bench: Benchmarking end-to-end speechllms in the wild.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Wildspeech-bench: Benchmarking end-to-end speechllms in the wild

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.331772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:710b58c2cd47bb8815a01a53bc11e7f60fc5abcbe90241d43bcf0fcb63f7a2f9

Observation 0d3ffd4d-214c-4e3f-b28b-b42b3dcfec6f · outbound

This paper cites MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.372663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:40b7d5ce026c1189bc94be614db1978b8b8c4c4b01e5b03b94313c86317f5d2d

Observation 7e80d426-1ee9-422c-b4d8-0e66d83a7bc0 · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.295224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:ad2f8acfdbb0d454512759e5c9482c5e7e515b334a3ced23c8c7ac4e53055864

Observation cebd606e-9085-4271-b8a9-81102bd57033 · outbound

This paper cites Multi-bench: A multi-turn interactive benchmark for assessing emotional intelligence ability of spoken dialogue models.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Multi-bench: A multi-turn interactive benchmark for assessing emotional intelligence ability of spoken dialogue models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.364368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:92d2fd04dac1d40731b96c9c68559f427e9de2c74d3efdda860e8d618c13b34c

Observation 44d56285-f5f5-4bc4-b632-72ac5a7c9ad5 · outbound

This paper cites Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.383937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:5bdee95a9b5d1d51890d23f973a06f77d00784a844bb77d2536c0e58c82f763f

Observation 432c3dd3-0ba3-4c31-b0da-e50fb7711d3c · outbound

This paper cites VoiceBench: Benchmarking LLM-Based Voice Assistants.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action VoiceBench: Benchmarking LLM-Based Voice Assistants

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.312480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:b2fa32c5a7a9feee9aef8c86f322dd7e115a50f4b823dd55333007fb274413cc

Observation 54f92202-0963-446c-b3c9-10a8abe0ae5d · outbound

This paper cites AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.327536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:fca2e1fcdd82bc854871934b461d9e5d6dd7c2e109fffac7d421c990c3b34abb

Observation 2ee8075f-3901-4ce0-9480-6a7e7e40ba33 · outbound

This paper cites SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.318295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:a013b8ccc69295cb5974371711bdc27c27fb8b9e64f5ce73b8cc97920baaa25d

Observation d981c1f0-b2bf-43c5-9c63-56797b1abc01 · outbound

This paper cites URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:34:57.301296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:b32019d7e236e4e8d517101d847136655381f507093ed6087b734511f30c7b79

Observation ca4b479f-2bd3-40ed-b410-323475f59dfc · outbound

This paper cites Vocalbench: Benchmarking the vocal conversational abilities for speech interaction models.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Vocalbench: Benchmarking the vocal conversational abilities for speech interaction models

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:34:57.343785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:9bf9f6761281542fdc1273ef2deaefa333d36b1e8dde39c5b153c8bb6c82fa16

Pith citing papers

Observation 1107d89f-ae49-4579-be5d-e44915fae6c3 · inbound

Audio Interaction Model cites this paper.

Audio Interaction Model DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-02T10:46:52.391705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T04:57:05.062465Z digest=sha256:5b5b5601afa424ea978c0493234faad5aec0ccc1da1c71b698f11033b8f7b87f

Observation 020238b3-4dcd-474d-b0b9-c218d45259c4 · inbound

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models cites this paper.

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:49:57.634773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T00:13:38.523397Z digest=sha256:f3d4a073090f9e5a08151d0a1a11de4e63cf401ecda41c6b34cbd6dd625abd97

Observation c1bd3ba1-4cf9-4664-ae10-c711551b2895 · inbound

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models cites this paper.

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-04T13:09:50.915862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T05:27:24.078053Z digest=sha256:c21b4d25a1de51a0fcac7fe952c174edf653af2b3d3a40b5d477a6d0ff16ed9e

Observation b7f1b8cd-a8a2-431f-af3a-160430d3f501 · inbound

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models cites this paper.

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-06-30T09:44:37.357107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T09:43:09.347118Z digest=sha256:17632ed31cbd613630d3464ffd9b00397a36402892ab043742ce738bed001282

Observation e5a96d3b-4301-4639-8848-e9d03eedb21f · inbound

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition cites this paper.

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:57.633052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:57.633052Z digest=sha256:6f29d826cb9233db06f070fbe7c3b7710c89dfbcee60031cf58433215e511c34