Pith. sign in

Paper Citation Record · LEDGER

WavChat: A Survey of Spoken Dialogue Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2411.13577.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13577 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:41:00.748722Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d76f053a-4f51-44c6-9c0d-6b8f7b739e43 · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey WavChat: A Survey of Spoken Dialogue Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:45:08.048139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:7bed2cb1dbd896de8b6ecb87697d2f0cb2ed494b41a1705b0ee91ee91fef56de

Observation fd7ba0be-36c5-4455-8098-6c4bf2e123b0 · inbound

PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs cites this paper.

PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs WavChat: A Survey of Spoken Dialogue Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:00.748722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:00.748722Z digest=sha256:f91542c36152bc80797a023a50ffc2a66f7543b9cec6a1bf9f89452f38ef8dd2

Observation 24297cbb-faa3-4c2a-a25a-9cf84fe8d9b4 · inbound

Speechless: Speech Instruction Training Without Speech for Low Resource Languages cites this paper.

Speechless: Speech Instruction Training Without Speech for Low Resource Languages WavChat: A Survey of Spoken Dialogue Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:15.917688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:15.917688Z digest=sha256:f44f89d45528d1da702fc166bb789daffe5179d36745d32e49e35eac790466c8

Observation 4c3dccc7-2e75-4285-ad77-48527129a280 · inbound

BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM cites this paper.

BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM WavChat: A Survey of Spoken Dialogue Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:09.865471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:09.865471Z digest=sha256:9f196963495adfad2c8988d7cdf05f509ce50da230156c3c422e1114d33f632f

Observation b3c742b7-aa8e-4ea6-b906-76cf35b79278 · inbound

Breaking the Barriers of Text-Hungry and Audio-Deficient AI cites this paper.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI WavChat: A Survey of Spoken Dialogue Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:51.286122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:51.286122Z digest=sha256:d67a6cdf96e81edeb176c7ea095b05035e2058c2eb2ffd0adea22b1d1e25bdde

Observation 44080555-c607-49e4-a760-a70df34c9de0 · inbound

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training cites this paper.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training WavChat: A Survey of Spoken Dialogue Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:11.320572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:11.320572Z digest=sha256:0827e12e6a86f3a943e2c9c8a6f203ac9c39bc8139a12700d1b29339556aa97f

Observation 681a5cdb-7ba7-4382-9e1f-b936d7d04130 · inbound

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model cites this paper.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model WavChat: A Survey of Spoken Dialogue Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.393147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.393147Z digest=sha256:ef1a9249bb3f0bd62e7ba66d2eb929b6bbe0da00294a1127fc9a6e3691bee6ed

Observation b65c5f37-2a5f-4cad-b67c-2cbb5e689769 · inbound

OpusLM: A Family of Open Unified Speech Language Models cites this paper.

OpusLM: A Family of Open Unified Speech Language Models WavChat: A Survey of Spoken Dialogue Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:36.084747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:36.084747Z digest=sha256:f31bcfe2e1b64fbb692b183a6941ad0b7b0479ce64ca3791cde23d19c88008d7

Observation ae0594c1-2ab7-40ed-84a5-211ae59b4860 · inbound

UniMind: Unleashing the Power of LLMs for Unified Multi-Task Brain Decoding cites this paper.

UniMind: Unleashing the Power of LLMs for Unified Multi-Task Brain Decoding WavChat: A Survey of Spoken Dialogue Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:47:10.364647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T07:43:26.736409Z digest=sha256:57a5d8df8681a998e4ab37399d401eeef4c862d44eaddb1b67f54054554abeb0

Observation 8c503e55-8ba0-4943-9f41-c006f5503240 · inbound

StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding cites this paper.

StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding WavChat: A Survey of Spoken Dialogue Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:33:09.245104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:33:09.245104Z digest=sha256:8c6c8680259c7d8185e3e8b26e064a2d901102f9369974467c3b85fc33f4becb

Observation f140ba98-4637-4a27-b434-b70b112dc254 · inbound

Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling cites this paper.

Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling WavChat: A Survey of Spoken Dialogue Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:11:24.285128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:11:24.285128Z digest=sha256:482922555c19047cb7642be4174d6e3fc468fe566018768fd8aff88e1433904e

Observation 74403b0a-948f-405d-9850-ed7df164adbb · inbound

Dual Information Speech Language Models for Emotional Conversations cites this paper.

Dual Information Speech Language Models for Emotional Conversations WavChat: A Survey of Spoken Dialogue Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:45:22.176873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:45:22.176873Z digest=sha256:1ed9bc0301b9cad48aff6397f718115722436c86110affc57bb0b2c9b90fdf93

Observation 5288e3c4-2220-4f74-bd8a-ccf2e787a69d · inbound

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs cites this paper.

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs WavChat: A Survey of Spoken Dialogue Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T16:46:49.184979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:46:49.184979Z digest=sha256:7483fe7ddc27b8a9012018424f5e01985977ca5fe06db9b6bd9a33ce799887af

Observation 9f744dbb-9fdf-44a8-9b65-209172234df5 · inbound

ChipChat: Low-Latency Cascaded Conversational Agent in MLX cites this paper.

ChipChat: Low-Latency Cascaded Conversational Agent in MLX WavChat: A Survey of Spoken Dialogue Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:44.007778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:44.007778Z digest=sha256:cb9d388769da2a96f4ffde60acd9786d0be0e409ba41e82123c8f95ffc6b0740

Observation ab9084d7-ed4c-4928-add3-05e6c2c8fd8c · inbound

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models cites this paper.

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models WavChat: A Survey of Spoken Dialogue Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:52:35.726992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T11:51:43.561210Z digest=sha256:4089df2bac9f3275e99b77dcb068839b89a4dafc6d27ee1894cd2172066e38a1

Observation b19cbf1b-6668-4905-b09c-0e80e2297a68 · inbound

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM cites this paper.

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM WavChat: A Survey of Spoken Dialogue Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T21:42:49.032029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:42:49.032029Z digest=sha256:c067825ca1f1cfc8664af09a136a77e4c67747285ca2ce4b599c431ee249cc2f

Observation 271fc78e-d208-4431-9e80-c67457bd918e · inbound

TiCo: Time-Controllable Spoken Dialogue Model cites this paper.

TiCo: Time-Controllable Spoken Dialogue Model WavChat: A Survey of Spoken Dialogue Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:39:35.866368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T00:38:52.182973Z digest=sha256:59c450c9512fe9fff01118b499c30421ae2f2e41566c41a750743d3c5433cd3c

Observation 535b2d08-65a5-4aa4-ba7a-7e720b089a69 · inbound

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff cites this paper.

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff WavChat: A Survey of Spoken Dialogue Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T20:09:57.922788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:09:57.922788Z digest=sha256:13f58728c77043fe6c9e0a0bbfbed187bc51a35c8596f0593833db910c126586

Observation 7bd73e9c-b565-485d-acf0-d87c0ab92ffc · inbound

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection cites this paper.

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection WavChat: A Survey of Spoken Dialogue Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:35:18.904751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T11:32:10.126062Z digest=sha256:0d7457eb7dbc5697370c3cf7fdb69bca7fbb7507d09c7b86cf7b47db35d05dc4

Observation 8ea7f99f-86af-441f-a610-0d9a707b88a3 · inbound

Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals unreliable Multi-Turn Behavior in LLMs cites this paper.

Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals unreliable Multi-Turn Behavior in LLMs WavChat: A Survey of Spoken Dialogue Models

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:22:21.217967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T02:18:07.162514Z digest=sha256:6263aa66b5dfe82c32550b06240f13264fae243b437d45755827dc6e776a7048

Observation e8014a72-e25e-48c7-9938-bfde1e7c5e27 · inbound

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models cites this paper.

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models WavChat: A Survey of Spoken Dialogue Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:41:26.363121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T14:29:18.348031Z digest=sha256:995badcc7ce34f0c8b4e2e9dcb0e9cb375741a64980c3635f694d21a3dfbc2a9

Observation 6de18e82-e85e-4747-9da7-4f588149be8f · inbound

VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models cites this paper.

VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models WavChat: A Survey of Spoken Dialogue Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:51:09.764621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T16:59:25.973854Z digest=sha256:91f5159ee55ac98a142b9f5c3fad38e35a5a33bf5b4dce8e71649383b0c308ba

Observation 36ee66e8-0806-429d-94e4-0367b3d78a91 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing WavChat: A Survey of Spoken Dialogue Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:50:55.760212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:adcf0dead030f75b8b9acd8472dfba03621e67815d4c27ea567d7cc316bf3515

Observation 7a2fed10-bd5c-417e-bed0-684bfb935262 · inbound

Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades cites this paper.

Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades WavChat: A Survey of Spoken Dialogue Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:03:17.758837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T13:02:53.960375Z digest=sha256:5a11694ed22b2a05ab19a7b28bd55f23d5d75e2f6c47d3ef5c28b9d5502842d8

Observation 6d257d6d-1135-4f57-aec4-5ffc20237160 · inbound

Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades cites this paper.

Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades WavChat: A Survey of Spoken Dialogue Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:15:00.392581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T19:12:32.672352Z digest=sha256:c5652f7cb880ed121a319482433152f94d634d08c8f8782c2c71ac0a4a519f93

Observation 3e1a5b93-42aa-42d3-b65a-0f686d7bb436 · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook WavChat: A Survey of Spoken Dialogue Models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:49.050610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:a8774a6269c01ca6d128a41ae72104664226ce4482a535050179684f922575a5

Observation 292c42a6-2bcc-4fdc-b4d3-5472e442d172 · inbound

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects cites this paper.

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects WavChat: A Survey of Spoken Dialogue Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:13.838719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T17:44:07.669223Z digest=sha256:54cfdf060ae802035ec2752962a4debe53a46e00fc8acb3d62c363b7d8cc66b1

Observation 56dd88e1-fd28-4878-aafe-63dad0995384 · inbound

UniVocal: Unified Speech-Singing Code-Switching Synthesis cites this paper.

UniVocal: Unified Speech-Singing Code-Switching Synthesis WavChat: A Survey of Spoken Dialogue Models

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:46:24.646473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T13:17:13.510587Z digest=sha256:8dd0d1a3b76f1c5da92d7177a5d85c27c2612f3e5de7f79a9ff868af998203b4

Observation 9e561681-2cfb-4359-be47-55f7f9d8f1e8 · inbound

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models cites this paper.

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models WavChat: A Survey of Spoken Dialogue Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:20:57.380720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T12:53:00.569655Z digest=sha256:13a267a116e466dc37429997030a93e59b2eefb71653f81a3e77cf80d54927a4

Observation 5615766a-21f1-4d73-8a28-d4a93a13076e · inbound

Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering cites this paper.

Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering WavChat: A Survey of Spoken Dialogue Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:39.504681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T13:22:52.428152Z digest=sha256:7f0d39f84b07059af92ff897696d69e55649fe9be6218b3f973b36462dffa1d0

Observation fc803355-0644-4879-9499-af55bdc620cc · inbound

PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue cites this paper.

PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue WavChat: A Survey of Spoken Dialogue Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:33.547578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T06:53:13.344795Z digest=sha256:3a889035dacd23270f1d5caeaf90891747a54308612cc3523dbcf0a9bac86830

Observation 0a6e3651-30ce-49cd-a0a5-a1d5c29ac73a · inbound

Integrating Facial Generation into Full-Duplex Spoken Dialogue Systems cites this paper.

Integrating Facial Generation into Full-Duplex Spoken Dialogue Systems WavChat: A Survey of Spoken Dialogue Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:41.319388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T12:08:55.070482Z digest=sha256:d6f8deccbf3c8c645ae1f3cb4b35e8c1077b10a22dc842ae634c81315faf8715

Observation 5bff76c7-2b3a-478d-bc3e-aaf2a390e528 · inbound

Audio Editing in the Era of Foundation Models: A Survey cites this paper.

Audio Editing in the Era of Foundation Models: A Survey WavChat: A Survey of Spoken Dialogue Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:09:49.441309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:74b180f44f1486cbecdde7128fd54dbecb5a303f4e950262ee343106432465dc

Observation c068f586-cd16-4bd7-9cdb-a91e167203ce · inbound

RedVox: Safety and Fairness Gaps in Speech Models Across Languages cites this paper.

RedVox: Safety and Fairness Gaps in Speech Models Across Languages WavChat: A Survey of Spoken Dialogue Models

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:52.893021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T04:37:00.399470Z digest=sha256:ce8ccf87a6b1fdeece9a1d9848a484ed48c6eccd2c16854079bb6c9ca0e1742c

Observation 16ed510f-853e-45a6-b971-8f702480c31a · inbound

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue cites this paper.

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue WavChat: A Survey of Spoken Dialogue Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.297693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T21:22:38.888528Z digest=sha256:487326818a10e64cf567688d2c6bfeedff97b456aa004a225dea5a5e5d055b60

Observation 1aacc01f-e37b-4fce-9c64-46d6f0f3761f · inbound

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue cites this paper.

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue WavChat: A Survey of Spoken Dialogue Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T08:53:13.408944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:53:13.408944Z digest=sha256:77cfb24d41f587a4f50dc885125d0190203d357bd3bbe2514d4fda840f7b33a8

Observation b65231ce-691f-49ab-bddc-1b16309664a0 · inbound

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs cites this paper.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs WavChat: A Survey of Spoken Dialogue Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.611257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:3a1fea309cb90ba8c22442461449bafae56a71f04b85233bdf4c400ddc0b5e6d

Observation 26ce0f7a-9521-40c9-a8ca-99e9b2e5484c · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment WavChat: A Survey of Spoken Dialogue Models

Reference 131

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.646123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.646123Z digest=sha256:a72e430b738bce29c6e583de7916ed5d2ffed7de6f096d6269ae18eb791beab5