Pith. sign in

Paper Citation Record · LEDGER

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine

As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2606.19453.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.19453 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T19:02:16.119373Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact24
  • verified fuzzy0
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dc8158c4-0919-40b5-bef4-628ca0407403 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Moshi: a speech-text foundation model for real-time dialogue

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:49:25.054154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:c6e5422c9c534955f2c0eabc8834b3dc19769090737bc3b67334eb421d64face

Observation 1c775997-0ae1-4ebb-bcca-133882e01dff · outbound

This paper cites MinMo: A Multimodal Large Language Model for Seamless Voice Interaction.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T02:49:24.994509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:1e607e6d7f0c9368bccb13b28dc47c0ea8418762d8b26c03aeabe0d3d48bb2f8

Observation 4b3cdabd-2207-42fd-842b-80b5c1a22b6f · outbound

This paper cites FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:25.022076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:c3601c8ecd319e87185f894effd4d3131efc249cf5abf48e8eb3f3b96617a89d

Observation 010a61b3-bee2-4c9f-b806-7a215fd51d02 · outbound

This paper cites From turn-taking to synchronous dialogue: A survey of full-duplex spoken language models.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine From turn-taking to synchronous dialogue: A survey of full-duplex spoken language models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:25.027785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:0a11f1c4f3e3249c61e70bf59c548a82b973a4fb47c6d867e9acd49bd8080640

Observation 799ea9b2-8e0c-43b9-b28b-7b8c2b3100eb · outbound

This paper cites FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T02:49:25.051912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:7543ec08cc9e19d5775eb208f98fcc6be94941bde073513ef15cc826f1efc549

Observation 681b7d68-a2dc-4532-ac1d-ede5959dabbc · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:25.016400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:a72f18467a33c96726fd0fa7b751d65f4873fc23921abad1e608f69fcb21826e

Observation 575e11c0-60aa-4308-9f3f-9cf019c65b49 · outbound

This paper cites Llama-omni: Seamless speech interaction with large language models.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Llama-omni: Seamless speech interaction with large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T19:02:16.119373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:8d3061e443e7da1bc7e39f86bec9a513d83b7bfcbb16d0b4ea7156d3fff54be6

Observation 5c3b866c-7d72-441f-8290-e74b830c22af · outbound

This paper cites Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T19:02:16.119373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:be3df8dcd1f44561fc8caa1d4d86b3cfe016842b8ce392bdff6889b7567dd32e

Observation 152022c0-615d-4c18-b4b9-b96624a0cb49 · outbound

This paper cites High Fidelity Neural Audio Compression.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine High Fidelity Neural Audio Compression

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:49:25.059455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:977052b54109db084de6939f258da2cd5cd6302115adf17fc580f61b5900f543

Observation 833063de-95b3-488c-a55a-f9a2e4206004 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:49:25.025115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:db538284e4f8066e6f1a4f37cece671efdbdc422e982f78d7dd4ae0d0a52aa00

Observation 92a57590-d5b1-4f10-865c-c281a3a039b6 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T02:49:25.010931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:d12eccce8bb22d4ea71100ebafbdb5ea47a2bc631c8dcbfab4b28dd0bd5d4fb9

Observation fe31b96f-3af2-446d-b4ae-f8375842a4b5 · outbound

This paper cites Wavtokenizer: an efficient acoustic discrete codec tokenizer for audio language modeling.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Wavtokenizer: an efficient acoustic discrete codec tokenizer for audio language modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T19:02:16.119373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:691109979b4fc580632df72ed330476843bee1ae123f35a1cf94df34db883574

Observation 4cf45c11-4d86-4d20-be97-ba8120d39acf · outbound

This paper cites Kimi-Audio Technical Report.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Kimi-Audio Technical Report

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:49:25.034679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:81fa834241e3eb7dc50c1b0c6e6611f79c98afc81e35b968de73efc2df8aeb77

Observation 20bb6061-74d9-4860-a7fd-357cfb726f17 · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:25.056791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:3e38c2d8137ab61353c3f83318c2b4dc5c8c52e78a6e4aa5c87cd3b73cc417e3

Observation 3b3933fe-3130-4293-b348-2dff0c8d9eea · outbound

This paper cites Fun-audio-chat technical report.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Fun-audio-chat technical report

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:25.049038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:7d6e3a815cc839406ade999a57c6dc9daff86dcd2b2fc4aa239c5fcab43fb4c3

Observation f2ad20fb-f375-4d63-a80f-a152b99208ad · outbound

This paper cites Covo-audio technical report.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Covo-audio technical report

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:25.024865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:db6f28bbe6bede8d7614ea0ccebdcfec95fa916cdb0e781856acc73ceec1c0ad

Observation 991f45d7-b0da-4287-9e3c-2d97b1d6cb92 · outbound

This paper cites Soulx-duplug: Plug-and-play streaming state prediction module for realtime full-duplex speech conver- sation.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Soulx-duplug: Plug-and-play streaming state prediction module for realtime full-duplex speech conver- sation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:25.019606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:626c28f7925dda2040199469cc0ac82a295c17383c24dd3bad9c642d4fd92592

Observation 77733285-e27c-4e3d-85f1-c568fd45cef3 · outbound

This paper cites FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:49:25.032467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:6d83bff53b0b26aae3d6b647c372831b629f02f96ebd4404b786c4cd6e5bba9d

Observation b3485cf5-69fe-4163-b873-5246e41cc53e · outbound

This paper cites Qwen2.5-Omni Technical Report.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Qwen2.5-Omni Technical Report

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:49:25.046220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:5610c63b4a781aaaea05c557d9b24abe529dc813a0edc02b4df15ab1ff4f33de

Observation 64a7b10f-e728-42bc-8c56-6c2e0ddecfdc · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:25.045148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:64c29e58a905b8eae84d8949df5052d89b0206e1b0aad45db8606f38dab0331e

Observation 64f10523-1aab-4cec-8513-bb4242b094ef · outbound

This paper cites Qwen2-Audio Technical Report.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Qwen2-Audio Technical Report

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:49:25.038040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:1d10d0f5fbaa098c760585e3484973d7d40efdb599ead955d251865d7e649995

Observation da34a01a-b062-45f0-a7d8-25496182aff1 · outbound

This paper cites Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:49:25.013734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:d5387b527e1d0740dbef6899bf5de202c73ad77a1ebe591ce436993a5d37da61

Observation 3a15380d-6b5b-4de3-a6aa-2d23e53fbbea · outbound

This paper cites Personaplex: V oice and role control for full duplex conversational speech models.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Personaplex: V oice and role control for full duplex conversational speech models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T02:49:24.987760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:5126a541a2d570e9e62d5a19b29503fc7d37cb599c80c0257569fd383860fe99

Observation 7d085be0-f106-4ec1-911e-a2b03394d8d7 · outbound

This paper cites MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:49:25.005579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:63e2137df1bd093f07f13420ca3cfcf278bbaf034d07090f22d16b81296203c5

Observation b8d0121b-1f37-4f4f-82bd-f1396cd54742 · outbound

This paper cites Easy turn: Integrating acoustic and linguistic modalities for robust turn-taking in full-duplex spoken dialogue systems.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Easy turn: Integrating acoustic and linguistic modalities for robust turn-taking in full-duplex spoken dialogue systems

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-26T19:02:16.119373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:1d8fff2bce4730f9b53dc2c3097f9378c07bf39840aeb288b6d88642d3fade83

Observation 5eae23d3-bf32-4797-8b48-f7fdbeb7024a · outbound

This paper cites Full-duplex-bench v1.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Full-duplex-bench v1

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-26T19:02:16.119373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:eac4fbab95e0d4ee4e4bd0b74f61cff35e5d67c70bc92e015e304889ff3b7f6e

Observation 2044f8b3-a0a2-4667-be8a-b1f4625482e3 · outbound

This paper cites Switchboard: Telephone speech corpus for research and development.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Switchboard: Telephone speech corpus for research and development

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-26T19:02:16.119373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:3cf76ed1f0cf77fb972a36aea712aaad36a3674e7d2cd9ad34dd9328a2a9f2d9

Observation 875a3707-031e-4019-bb9f-7eb40f23444c · outbound

This paper cites Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:25.027416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:540d6ed0242bbfc72e2b375594af6cf97d900ca60bf3eefde3ca4f027c4db5cc

Observation 12114a70-a4af-4716-abd7-3990d52f97bd · outbound

This paper cites HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine HumDial-EIBench: A Human-Recorded Multi-Turn Emotional Intelligence Benchmark for Audio Language Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:49:25.022295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:cfb787f5bc10c7e4d0309d56508919628f9c2bee70b967df0d103deae679e456

Observation df159f2a-4ea6-4f3b-9393-25619b0c9f7a · outbound

This paper cites Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Dual-Axis Generative Reward Model Toward Semantic and Turn-taking Robustness in Interactive Spoken Dialogue Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:49:24.999926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:bd676d29e09e2f95b8742b5ee8c8ae56c700aea01525bc7a5e1bb625a2d13bd4

Observation 94e724fe-7e45-4836-9408-5f39b7802450 · outbound

This paper cites Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T02:49:25.047644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:789a88c06db9b569a3269c13e5186b1f2900bfacbce67871e1afe44d17b29f66

Observation 0e1281f3-4fb8-42f9-8731-b258e47814ed · outbound

This paper cites Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:25.019234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:901a3b2fe8d06751829bb93cb527d21cd2896c994456407535b8b50dbe525808

Observation f1ae03e7-9214-4a26-9416-82a418665422 · outbound

This paper cites MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine MTR-DuplexBench: Towards a Comprehensive Evaluation of Multi-Round Conversations for Full-Duplex Speech Language Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:49:25.030118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:105d5c064d68d43621c5b9c18fde4ad86c3d33619a4e2a83a7037a2d0911761c

Observation caff9bb2-9bed-430d-b66a-ac9fa46131f1 · outbound

This paper cites Flow Matching for Generative Modeling.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Flow Matching for Generative Modeling

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:49:25.036883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:6428646bdfdd1e113321f86a26e649e3fde587dc9b4a51af921349ef66f364be

Observation b61fcfa9-fe67-4a94-80c4-605cff841752 · outbound

This paper cites Revisiting Feature Prediction for Learning Visual Representations from Video.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Revisiting Feature Prediction for Learning Visual Representations from Video

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:49:25.044035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:9facd653f07969c829e2c0b88068bccbd75ceefbcedcd9ea626e1b27c6b6fa2d

Pith citing papers

No inbound Pith citation observations are available.