Pith. sign in

Paper Citation Record · LEDGER

WavChat: A Survey of Spoken Dialogue Models

As of 23 August 2026, this Paper Citation Record lists 100 of 258 outbound references and 47 inbound Pith citation observations for arXiv:2411.13577.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13577 v2

Coverage vector

measured 100 of 258 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:13:57.494016Z

measured 147 of 147 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 47 of 47 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:52:07.865325Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 258 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved100
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 9c26b626-7630-461b-a1f5-3fe009b28cd0 · outbound

This paper cites GPT-4 Technical Report.

WavChat: A Survey of Spoken Dialogue Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:56.957483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:56.957483Z digest=sha256:a59c89a557471c4383b788e20fe89417fdc922a0ee974d932950ec981b7ba855

Observation e9c2d8f8-9a71-4eab-b4b4-1096737dc39c · outbound

This paper cites MusicLM: Generating Music From Text.

WavChat: A Survey of Spoken Dialogue Models MusicLM: Generating Music From Text

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:56.962819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:56.962819Z digest=sha256:99cd6a107265c5281159f282f2e057a24b9ee7417d2ba8934fd0b1be58af1d6e

Observation 4ecaa38d-9edb-4a0e-9628-3526b7d01beb · outbound

This paper cites HILCodec: High-Fidelity and Lightweight Neural Audio Codec.

WavChat: A Survey of Spoken Dialogue Models HILCodec: High-Fidelity and Lightweight Neural Audio Codec

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:56.968238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:56.968238Z digest=sha256:c40695b03a333d42c764cf069a0ee655c3e664fb12d2e7773b04511b92c2ea71

Observation 002ca7ce-610f-4862-a914-196ffd609c03 · outbound

This paper cites APCodec: A Neural Audio Codec with Parallel Amplitude and Phase Spectrum Encoding and Decoding.

WavChat: A Survey of Spoken Dialogue Models APCodec: A Neural Audio Codec with Parallel Amplitude and Phase Spectrum Encoding and Decoding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:56.974234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:56.974234Z digest=sha256:29a211489cad688cc6da87883e36514a58503bf57c3de74928bb3d0d53180551

Observation ea9d8ba8-8bc2-43ed-b261-2854fb0377fc · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

WavChat: A Survey of Spoken Dialogue Models Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:56.979573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:56.979573Z digest=sha256:c2206263b91c51766f3970bbb6caa0e134fb2022a213543e4394e2f49eaa95dc

Observation 2d0a9442-fe6d-4354-83fb-95ad0a456654 · outbound

This paper cites PaLM 2 Technical Report.

WavChat: A Survey of Spoken Dialogue Models PaLM 2 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:56.985115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:56.985115Z digest=sha256:c35fba49cd78f4ae37e19b6d3a2b421cd15cb051f9e9378da70806c139d106e2

Observation b7b9cc61-d310-4e55-9554-f66189e4215b · outbound

This paper cites SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words.

WavChat: A Survey of Spoken Dialogue Models SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:56.991259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:56.991259Z digest=sha256:21edead34618fa27960bdc4df66b01a500dae8a6473304704a60b3bcb6ec60f6

Observation 956ed686-9594-4924-b517-422a317da9f2 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

WavChat: A Survey of Spoken Dialogue Models Common Voice: A Massively-Multilingual Speech Corpus

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:56.996509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:56.996509Z digest=sha256:765e8d44bd52f096ff04efc47027a5a410b1f4f720af29cc0c5736af099c2616

Observation a393ab7b-61a5-490b-98a7-7beda0ca40bf · outbound

This paper cites XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale.

WavChat: A Survey of Spoken Dialogue Models XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.001454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.001454Z digest=sha256:0e7a542f94d06ac96f07fca5844fd31c077e334f8522593e46c998771a3d9657

Observation 3e28eb88-9e4c-4b9f-9016-4a6e2e8443a8 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

WavChat: A Survey of Spoken Dialogue Models wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.007075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.007075Z digest=sha256:5c07fadede5225f22d862e3fab79e242a71325c376a55fbf8e1da9635d8d3137

Observation 315dc7ec-e5eb-4f92-b4a9-501bac386e64 · outbound

This paper cites Qwen Technical Report.

WavChat: A Survey of Spoken Dialogue Models Qwen Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.012278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.012278Z digest=sha256:cebcbb4fb2170d84ede3f800238ff0d36075227f3e28f021a7969b2c81ed17f7

Observation b4a12570-d991-47bd-bf60-b71f05c3238b · outbound

This paper cites An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling.

WavChat: A Survey of Spoken Dialogue Models An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.017135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.017135Z digest=sha256:5734d6e2e186da5fba5b9211bffa57df57e25375c9cf1112a4a4277ba62ddeb6

Observation 76dcd302-7a21-4582-a493-5dc93222cddd · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments.

WavChat: A Survey of Spoken Dialogue Models Meteor: An automatic metric for mt evaluation with improved correlation with human judgments

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.021600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.021600Z digest=sha256:79f47e2a77439062e482acfd433777c6a3d7026f01fdfc382f246177c751db24

Observation e6800971-c408-481c-8671-963e5dd4fe91 · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

WavChat: A Survey of Spoken Dialogue Models Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.026191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.026191Z digest=sha256:f8e72c2f54a978c0c7d25c7e77e3accaef5f5fcd0abd268d713b99f222cd3183

Observation 1daf4044-5c06-4330-8d8f-c482137b6063 · outbound

This paper cites Curriculum learning.

WavChat: A Survey of Spoken Dialogue Models Curriculum learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.030773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.030773Z digest=sha256:27c1ef40b80715b6ef61672023b1d183c075778ebe4a44f22dae209526b7afb3

Observation dab89427-e657-4506-9ccc-2cebbf9860b7 · outbound

This paper cites Medleydb: A multitrack dataset for annotation-intensive mir research.

WavChat: A Survey of Spoken Dialogue Models Medleydb: A multitrack dataset for annotation-intensive mir research

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.035011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.035011Z digest=sha256:9c4c8739db50e18634060e541520dfadf458da8a5d50f4645f1a29fdfe16c559

Observation d2af126a-f211-4433-8a83-ed71a03ba5d4 · outbound

This paper cites A comparison of sound segregation techniques for predominant instrument recognition in musical audio signals.

WavChat: A Survey of Spoken Dialogue Models A comparison of sound segregation techniques for predominant instrument recognition in musical audio signals

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.039424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.039424Z digest=sha256:749dca7ce5d6138b0d868344245551299aab3b98768ec2ca89e976e177420c63

Observation 4765c84f-026a-49d4-966a-021719ab9b7e · outbound

This paper cites Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline.

WavChat: A Survey of Spoken Dialogue Models Aishell-1: An open-source mandarin speech corpus and a speech recognition baseline

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.043938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.043938Z digest=sha256:16a7ce218640d58de1aa2f696e27a950bb98d76f3bbaae25957671ce4cb69e1f

Observation 258ccbfc-8bbf-4433-b616-92d0b2f50e44 · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database.

WavChat: A Survey of Spoken Dialogue Models Iemocap: Interactive emotional dyadic motion capture database

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.048395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.048395Z digest=sha256:e5f998055d6c92e5a7d0d4572318f9227616aee4b783ccd058bfd28c8ef0a0ea

Observation d2fb2d2e-336a-42d4-bf86-be0f562b8128 · outbound

This paper cites Msp-improv: An acted corpus of dyadic interactions to study emotion perception.

WavChat: A Survey of Spoken Dialogue Models Msp-improv: An acted corpus of dyadic interactions to study emotion perception

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.053955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.053955Z digest=sha256:c855ba89d15553015c95960b39b10faa5f8c9160a2b865057c722469ffccd164

Observation 3b44d82e-5b4b-41e4-b5dc-22689577b166 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

WavChat: A Survey of Spoken Dialogue Models XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.058757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.058757Z digest=sha256:4ffd31b1d44ae6f98801b1fbf1c6aed1517d2c90542059bb6cb72f165d16529c

Observation 018c40ae-3892-4bec-b1e2-e4c4d8558bbc · outbound

This paper cites Mistral 7B.

WavChat: A Survey of Spoken Dialogue Models Mistral 7B

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.063979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.063979Z digest=sha256:a970be8c48147371c683b4855d4e554365080f58c8de8384584bbba425dd60b8

Observation ceeb780d-66d9-4dca-a1e0-b4cfcddc395c · outbound

This paper cites Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback.

WavChat: A Survey of Spoken Dialogue Models Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.069312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.069312Z digest=sha256:130894753420cee529f499d99d2cb88d3e09dd3d0e1b17297b111c6a7ed4df06

Observation 4db6cd1f-8b9a-4688-a0af-bcaaf51a1b75 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

WavChat: A Survey of Spoken Dialogue Models GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.074793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.074793Z digest=sha256:6ea6b50f87937dd8125a93758fb8ed23322c8dc9da587cb810df09d2347e1e77

Observation c02b3be4-debd-4f67-8ee1-6d6aa4965463 · outbound

This paper cites EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions.

WavChat: A Survey of Spoken Dialogue Models EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.080098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.080098Z digest=sha256:eaa65cb7134981c8040b7743e417f2fa209720fe7ca50b67c3c308dbc05e43e6

Observation 3f7b9dba-156d-462e-9bf3-948ea0e899da · outbound

This paper cites Evaluating Large Language Models Trained on Code.

WavChat: A Survey of Spoken Dialogue Models Evaluating Large Language Models Trained on Code

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.085035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.085035Z digest=sha256:eba273cb14faa97d4ede1e6142399869404a243f145dfeae704dcdfbaf84a308

Observation c2a1efb3-59e4-437e-ae96-2947bae9c2b3 · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing.

WavChat: A Survey of Spoken Dialogue Models Wavlm: Large-scale self-supervised pre-training for full stack speech processing

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.090451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.090451Z digest=sha256:9b5fbae12bc53dbf1f52fb94d087afefc74bab8ff016c8b30847377632e59132

Observation 7df371f9-bab8-4f0b-931c-14d381ef70e5 · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

WavChat: A Survey of Spoken Dialogue Models BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.096118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.096118Z digest=sha256:d35bc5ea152354cf290b9d927d9831045b04098f4522ceafe42ae6168352b014

Observation 5b47d8fc-ad34-4938-87f1-403ed37ecbb8 · outbound

This paper cites VoiceBench: Benchmarking LLM-Based Voice Assistants.

WavChat: A Survey of Spoken Dialogue Models VoiceBench: Benchmarking LLM-Based Voice Assistants

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.101245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.101245Z digest=sha256:2c090e046bbfe9bd56b5b1ee46eb746e5a44d9f927d0e379c05a2b39c4c97777

Observation 8c645ca3-7c0a-47bc-a393-ef41920af236 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

WavChat: A Survey of Spoken Dialogue Models F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.106570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.106570Z digest=sha256:0f7bf50fac9406e7d5425d4817fcc47820d09b0470d0401651f6851fa2d725ad

Observation ef5c41f2-bef0-43b9-8800-83c3b39d8a6a · outbound

This paper cites Audio albert: A lite bert for self-supervised learning of audio representation.

WavChat: A Survey of Spoken Dialogue Models Audio albert: A lite bert for self-supervised learning of audio representation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.110882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.110882Z digest=sha256:5b3be4b674794ad6a0a1b19ad72be6c8fc7f721920a1e1454ddf81b80f608fd6

Observation cac48763-3162-485f-909a-b2c4aed89fd4 · outbound

This paper cites Coding Speech through Vocal Tract Kinematics.

WavChat: A Survey of Spoken Dialogue Models Coding Speech through Vocal Tract Kinematics

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.114967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.114967Z digest=sha256:63f231f9d0d3b929d192dc9c5f25814d535bae04cbf54ea945274ba517822590

Observation c593901c-4945-49fb-aee4-e89607ae5845 · outbound

This paper cites Qwen2-Audio Technical Report.

WavChat: A Survey of Spoken Dialogue Models Qwen2-Audio Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.120817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.120817Z digest=sha256:d34f34e5b3fd1e35e2587c086a837f6eadfb9c2c2929071c4bd566ed427dfe47

Observation cf343a91-b7a9-4c22-ad3d-d1e7b00db920 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

WavChat: A Survey of Spoken Dialogue Models Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.125555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.125555Z digest=sha256:079aa9444ed43ca1191174c4e3072591094c253f3c03b55c5d400f9004684bdc

Observation be047d4b-f468-479f-84da-bd1ead5288e8 · outbound

This paper cites Vector-Quantized Autoregressive Predictive Coding.

WavChat: A Survey of Spoken Dialogue Models Vector-Quantized Autoregressive Predictive Coding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.130253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.130253Z digest=sha256:b5c22eb4a0c8aecc85c01ebdb4f6e543b11105e92100f0b12d8dd32e4826ba1b

Observation ed8bf43c-20bd-4ecf-a844-feff12ec830e · outbound

This paper cites MusicRL: Aligning Music Generation to Human Preferences.

WavChat: A Survey of Spoken Dialogue Models MusicRL: Aligning Music Generation to Human Preferences

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.135983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.135983Z digest=sha256:7aa0e9f7816578899b2803e1b5e4d9eb8c3c97fec12e74a8c9732b82f43f53be

Observation ce81594f-53ac-49bc-acf0-c7e9a3e76d4c · outbound

This paper cites The fisher corpus: A resource for the next generations of speech-to-text.

WavChat: A Survey of Spoken Dialogue Models The fisher corpus: A resource for the next generations of speech-to-text

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.141036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.141036Z digest=sha256:a000949c2b4b1ad7ec5b688886982eeed23aee3d39be35b8b9d72f8e98145340

Observation c1da52ec-8c73-41a9-868d-68f46c881fcc · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

WavChat: A Survey of Spoken Dialogue Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.145838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.145838Z digest=sha256:58a2bd658f2bf951d95e1fa279be57f3245e6359cf851685ab053af024611ead

Observation c28f2a5d-7bd0-4402-ac24-8f356ba2676b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

WavChat: A Survey of Spoken Dialogue Models Training Verifiers to Solve Math Word Problems

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.150893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.150893Z digest=sha256:a4689eca0a5bc09180132a9a1c68de559c05045b6763bdfd56b6b602aed0c1a3

Observation 62b36e97-f652-4c13-bb1d-f662964f8706 · outbound

This paper cites Simple and controllable music generation.

WavChat: A Survey of Spoken Dialogue Models Simple and controllable music generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.155813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.155813Z digest=sha256:84a5f73582c5c6638be976de2cbc6fcaf8ff7f1536b223dc0deec39f5d908f46

Observation 4d1dd1b5-327b-4a08-8f2b-1ff04dd74303 · outbound

This paper cites SpeechVerse: A Large-scale Generalizable Audio Language Model.

WavChat: A Survey of Spoken Dialogue Models SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.160614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.160614Z digest=sha256:6fae6aad285de294bb6564b561b8ef8e335e49e62daa086ad08b0907cb07a110

Observation f6ec07a2-e710-4c85-902a-ca44a7fbba2b · outbound

This paper cites FMA: A dataset for music analysis.

WavChat: A Survey of Spoken Dialogue Models FMA: A dataset for music analysis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.166013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.166013Z digest=sha256:4ba2695ee9f1130672e424583fb51152112aa2a2eadb838c91d581c055e9026b

Observation a93b1fac-b5e3-48d3-8260-ea26665cacb8 · outbound

This paper cites High Fidelity Neural Audio Compression.

WavChat: A Survey of Spoken Dialogue Models High Fidelity Neural Audio Compression

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.171080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.171080Z digest=sha256:eccdcc57ebcb911e7f1d0c87b12deaa3649186a302e813697f03eb3fb36cc600

Observation 29fb114e-c123-4060-8e32-bf2cf5c3210b · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

WavChat: A Survey of Spoken Dialogue Models Moshi: a speech-text foundation model for real-time dialogue

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.176108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.176108Z digest=sha256:2d87b905cd826e0195f15f313a0d697c1408e3aaab6a8def030b033c9792bfb5

Observation 0cefef90-71e3-40e3-a52f-0be6618637f6 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

WavChat: A Survey of Spoken Dialogue Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.181222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.181222Z digest=sha256:4bae4b294a297a3153376626729fdc4841441667af2540c9e6eb22afc6f002cb

Observation 22a5a61f-24bb-4d91-9549-a875aab4479e · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

WavChat: A Survey of Spoken Dialogue Models Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.186385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.186385Z digest=sha256:80d617bb460a423ed52839be4d87c96681d4d266c31c89f38d82d69dc62c966a

Observation 2b4d91b7-1791-4e0c-8b0f-a307743f6348 · outbound

This paper cites Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey.

WavChat: A Survey of Spoken Dialogue Models Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.190980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.190980Z digest=sha256:8cbd0d8aa23f358838a3d95d59732f54b138b50aebfc04762e4aec18f7f79190

Observation 30ba73b8-d666-4736-81fb-0eb8a5cd1376 · outbound

This paper cites AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale.

WavChat: A Survey of Spoken Dialogue Models AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.195498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.195498Z digest=sha256:03cd3b14dc2e33d1608a9c3072c0c3076438af198a7e654cc88fa280aedb344e

Observation f9ce1a66-22c5-4ec6-b4da-662fd9ea569f · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

WavChat: A Survey of Spoken Dialogue Models CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.200756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.200756Z digest=sha256:37cbf145bd0888cd33feafe53cf025cd5d1a7f31b7408a4df751fecf11376ca5

Observation 5a316dd0-6089-4e12-b772-fe85c4600f4e · outbound

This paper cites LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT.

WavChat: A Survey of Spoken Dialogue Models LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.204938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.204938Z digest=sha256:d613a5575596878cc1d736705051d30c4de418890ef6c263c7a99b912f5d807c

Observation 2f229422-b4d8-4554-8113-4f578c56de24 · outbound

This paper cites Funcodec: A fundamental, reproducible and integrable open-source toolkit for neural speech codec.

WavChat: A Survey of Spoken Dialogue Models Funcodec: A fundamental, reproducible and integrable open-source toolkit for neural speech codec

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.209731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.209731Z digest=sha256:3aff37bd22777217210ef3405f0c8acdd52f25ddbf0002e076f66eb2d8409ef0

Observation bb83fbb4-cbb0-442d-94fa-42c7379c0d8b · outbound

This paper cites The Llama 3 Herd of Models.

WavChat: A Survey of Spoken Dialogue Models The Llama 3 Herd of Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.214193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.214193Z digest=sha256:28301ba8ff9e4890bd81a25bf3f64d0f2a8c96407ee5d40798e0c49827f4b926

Observation 40b767ac-f322-4d9e-983e-73d4f7a3b8d4 · outbound

This paper cites Some signals and rules for taking speaking turns in conversations.

WavChat: A Survey of Spoken Dialogue Models Some signals and rules for taking speaking turns in conversations

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.219888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.219888Z digest=sha256:23213302179ab3779f8a62c919b76d6195655bc1cb59d1cc0976fb7c63b72f17

Observation fcc02c13-c6d9-4d37-969e-037da153b3b7 · outbound

This paper cites On signalling that it’s your turn to speak.

WavChat: A Survey of Spoken Dialogue Models On signalling that it’s your turn to speak

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.224703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.224703Z digest=sha256:cce3993e44bc9073666ca662912032cdf2e84e5f7a883fb522f46d06b152e74c

Observation b49f78fa-0bd8-428e-9379-b744193cd15c · outbound

This paper cites TurnGPT: a Transformer-based Language Model for Predicting Turn-taking in Spoken Dialog.

WavChat: A Survey of Spoken Dialogue Models TurnGPT: a Transformer-based Language Model for Predicting Turn-taking in Spoken Dialog

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.229795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.229795Z digest=sha256:1ae5238a2d4cd9e78c968ed6a95387cd79bc6510017090e081a4d551eeea14b5

Observation 4670fd8a-6965-4992-afbc-6b8c907bc4e8 · outbound

This paper cites Neural audio synthesis of musical notes with wavenet autoencoders.

WavChat: A Survey of Spoken Dialogue Models Neural audio synthesis of musical notes with wavenet autoencoders

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.235031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.235031Z digest=sha256:15e91f252a96c2f96936fd633b08533013144c366f64688be9de8d4b120463a5

Observation c7753991-651d-4217-b129-458db8c2962a · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

WavChat: A Survey of Spoken Dialogue Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.240429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.240429Z digest=sha256:8b96dd25441bb67a211e4b56d87b00a6fc2d5581ece66293502dbe71c67071ee

Observation b7443e8b-4bfa-4c66-a79e-5940a9a44655 · outbound

This paper cites MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain Conversation.

WavChat: A Survey of Spoken Dialogue Models MMDialog: A Large-scale Multi-turn Dialogue Dataset Towards Multi-modal Open-domain Conversation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.246032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.246032Z digest=sha256:6f925b9f447c887ec63721dc65104ac33334c035d6d80f161b1c349a660b4dd6

Observation c892ff8d-7257-4dde-9b61-a0165ba24c30 · outbound

This paper cites Meisd: A mul- timodal multi-label emotion, intensity and sentiment dialogue dataset for emotion recognition and sentiment analysis in conversations.

WavChat: A Survey of Spoken Dialogue Models Meisd: A mul- timodal multi-label emotion, intensity and sentiment dialogue dataset for emotion recognition and sentiment analysis in conversations

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.252598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.252598Z digest=sha256:e5603c8160e9d2d0fdee6250e2cdf9c04fdb7f5aef5aeb01ae63bf92d845bccc

Observation 7fbd5429-1d82-4448-8d25-882be6851d98 · outbound

This paper cites Fsd50k: an open dataset of human-labeled sound events.

WavChat: A Survey of Spoken Dialogue Models Fsd50k: an open dataset of human-labeled sound events

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.258464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.258464Z digest=sha256:401e4df12d4a494180c93dca2a7ac088dc43c6d0467a80aa33df217106d5ca3a

Observation a3d0aa0f-6dfe-4ed6-b2a5-2b9f75f34491 · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

WavChat: A Survey of Spoken Dialogue Models VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.264011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.264011Z digest=sha256:3aeda5bd9d0fe1840f0a3f0bc299db2ecf4171025e24b4064f32f6bb7cd74fce

Observation 320fcf89-fc93-43c6-9d0a-23b137322614 · outbound

This paper cites A new algorithm for data compression.

WavChat: A Survey of Spoken Dialogue Models A new algorithm for data compression

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.269964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.269964Z digest=sha256:9e8acb86441e4c34b086b7421fee6b6cfa173d9bdcfaabaa5634fdc03605aaa1

Observation 424a669b-2016-4d62-a083-84d59e362c72 · outbound

This paper cites The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage.

WavChat: A Survey of Spoken Dialogue Models The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.274992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.274992Z digest=sha256:6105b17224be110b54c2087fb620e09cfcf5c57dd6d71edf1a8fb168761f0556

Observation b2eef2e6-8a45-4db4-ac5d-13225d60b5d2 · outbound

This paper cites Augmentation Invariant Discrete Representation for Generative Spoken Language Modeling.

WavChat: A Survey of Spoken Dialogue Models Augmentation Invariant Discrete Representation for Generative Spoken Language Modeling

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.280711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.280711Z digest=sha256:39bd782b829d7cd7fcad6074d39f25250b975e935aaa334f2a2225f8cf5959a2

Observation ef5af30d-ea6d-4a93-8541-893a3a1f877d · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

WavChat: A Survey of Spoken Dialogue Models Audio set: An ontology and human-labeled dataset for audio events

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.285812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.285812Z digest=sha256:4233acafe4f94491231bc4019cf24990573af78b5471cecba5a91d6a99601a68

Observation cb5e90a7-14b3-4374-b708-ab3f1a2d3636 · outbound

This paper cites Audio Dialogues: Dialogues dataset for audio and music understanding.

WavChat: A Survey of Spoken Dialogue Models Audio Dialogues: Dialogues dataset for audio and music understanding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.290608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.290608Z digest=sha256:515db6980b9893e1b8a03a31363985489cc834b13d7fd5ff6544a1419c7f4d93

Observation f44c55d5-6058-46b7-be64-741411121f5b · outbound

This paper cites Joint audio and speech understanding.

WavChat: A Survey of Spoken Dialogue Models Joint audio and speech understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.295612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.295612Z digest=sha256:86cc4cab96f3b4e2b39b1d7d1fa6e6739141cf8506d1ccfedb705ad3b5e86c71

Observation 06bc3aeb-6d83-4f5a-bdd0-4ea1250cd62b · outbound

This paper cites Listen, Think, and Understand.

WavChat: A Survey of Spoken Dialogue Models Listen, Think, and Understand

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.300573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.300573Z digest=sha256:e7ad0e86183852fc2c86e9fdda2c14554e66e50e84a424bced9f20e4cb2f00b1

Observation 3e5d15ca-1908-4d76-80bd-966b6e98ba46 · outbound

This paper cites Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks.

WavChat: A Survey of Spoken Dialogue Models Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.305300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.305300Z digest=sha256:adc8c0d2bf375d569b3b7f1e2d57ac8fa4e992b13a29c830a90ff26faf7f71a9

Observation 65cf570b-5c58-4d67-86c8-06a44039bacd · outbound

This paper cites SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis.

WavChat: A Survey of Spoken Dialogue Models SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.309591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.309591Z digest=sha256:36ec11d7bb203a0b647e3c14fc3c83653cad067647c3e25fdc1669174ff8a201

Observation d9d32c8f-cb84-44a1-9072-da299d4e3e45 · outbound

This paper cites Prompttts: Controllable text-to-speech with text descriptions.

WavChat: A Survey of Spoken Dialogue Models Prompttts: Controllable text-to-speech with text descriptions

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.314935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.314935Z digest=sha256:e3bdd80c32ac930e558ed1845d8bb2366647fcd000f5133f29e5c70bed6dff8e

Observation 2b7b6ff2-74d1-46e2-a9fa-00471ee69a4f · outbound

This paper cites Prediction of turn-taking using multitask learning with prediction of backchannels and fillers.

WavChat: A Survey of Spoken Dialogue Models Prediction of turn-taking using multitask learning with prediction of backchannels and fillers

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.320867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.320867Z digest=sha256:b1a47327c317b622868ecabc2cadcd227ba0103c6f7ce6cc3087f2a0bd4be4d7

Observation 9737d212-b214-4e07-ad82-87b2e567048c · outbound

This paper cites Turn-taking prediction based on detection of transition relevance place.

WavChat: A Survey of Spoken Dialogue Models Turn-taking prediction based on detection of transition relevance place

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.327107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.327107Z digest=sha256:341b2f26d0a900009504a771672e8dcae5edfef003bd4fff99399dc38083dc40

Observation c2d1c409-67bf-4a72-a12b-104db40e5015 · outbound

This paper cites Textually pretrained speech language models.

WavChat: A Survey of Spoken Dialogue Models Textually pretrained speech language models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.332330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.332330Z digest=sha256:814a127de09f50824b679868ff2ca0e3729166fe86157c1eb60ab0ba7c0bef3e

Observation 3d7304d9-8e79-476a-8f6e-2b93f738b65a · outbound

This paper cites Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation.

WavChat: A Survey of Spoken Dialogue Models Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.338204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.338204Z digest=sha256:69f0ab1fd6906c726ccd14ee4f5e02c936599c5baa396c3a9209d2921fb7f14f

Observation f45b7f13-d2c4-477a-bccc-97b8b805163f · outbound

This paper cites Measuring Massive Multitask Language Understanding.

WavChat: A Survey of Spoken Dialogue Models Measuring Massive Multitask Language Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.343744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.343744Z digest=sha256:09296b9fb1ac2bb3f2dcfeef3dd6fdfad66685ac86a7ebc99e82e81516405494

Observation d8c9351f-eaa2-49e5-8f47-6030e0f2cbca · outbound

This paper cites Cnn architectures for large-scale audio classification.

WavChat: A Survey of Spoken Dialogue Models Cnn architectures for large-scale audio classification

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.349347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.349347Z digest=sha256:e3cd2398d4e4ffd32dc00903a103c916ec331b99539bb680ce8a2208981209c0

Observation dcebdad7-c22c-42f3-b06d-3063928ef62e · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

WavChat: A Survey of Spoken Dialogue Models CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.355821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.355821Z digest=sha256:ed919665aa4f92d27348fdb73ad1e45172a5312e0fa785f055c7908616191bed

Observation 4983beac-99aa-4b8f-be82-d5ce90e852b9 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

WavChat: A Survey of Spoken Dialogue Models Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.363150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.363150Z digest=sha256:d93e81aaac3e2255a13503a8644ae6a3e68379a5f901dc8b55045ec531989618

Observation c156cbbf-a0d3-4c34-8ea9-10c02d575d6b · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

WavChat: A Survey of Spoken Dialogue Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.369774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.369774Z digest=sha256:46ee2c50c073b49ddade04f0e23299c01fa6f6b553a5d30018a525f2ae407780

Observation 800bb696-9382-4988-94b7-9b2b9178d651 · outbound

This paper cites WavLLM: Towards Robust and Adaptive Speech Large Language Model.

WavChat: A Survey of Spoken Dialogue Models WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.375568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.375568Z digest=sha256:5805765e82a5ae70926e226079abdf29c15681d7e2a38e3e899649009118ee96

Observation e68359b0-08d7-421c-ac43-232da744a15f · outbound

This paper cites Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation.

WavChat: A Survey of Spoken Dialogue Models Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.381615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.381615Z digest=sha256:23f6a65c5246f8280f5b06c8cec26ec4438a6f69aaac9fac19a57e8d1d98b1c6

Observation 78790e59-b9bc-4cfd-9456-50750a7d4d98 · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

WavChat: A Survey of Spoken Dialogue Models A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.389289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.389289Z digest=sha256:3ad6325d065bd771131b28d6803acf67edc768b79217aa593f3ae7cde46fa7c4

Observation 74190d8c-0307-4a0a-bc1f-7a279eae948e · outbound

This paper cites Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models.

WavChat: A Survey of Spoken Dialogue Models Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.394628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.394628Z digest=sha256:5010f377d42410a6ce58cd401144c090fd3dae25b39a521e1bfa6f77f1919e51

Observation 9636719a-5540-4c0f-b20a-eea6761db9b7 · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head.

WavChat: A Survey of Spoken Dialogue Models Audiogpt: Understanding and generating speech, music, sound, and talking head

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.404916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.404916Z digest=sha256:9fd37dbae38fc8f4fcdb18b144ccc2a1719e2e3839b52ad99fd281b04e36fbd4

Observation 8fac6b86-0a5d-4cc2-a969-cf23b331ccaa · outbound

This paper cites SPIRAL: Self-supervised Perturbation-Invariant Representation Learning for Speech Pre-Training.

WavChat: A Survey of Spoken Dialogue Models SPIRAL: Self-supervised Perturbation-Invariant Representation Learning for Speech Pre-Training

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.409432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.409432Z digest=sha256:f91bd3af73c7fa4ff6dc50333283ded9446f74117f8c6642ecd08d0570a03d68

Observation b5802cc6-2358-49c6-9ae1-f78fe999229d · outbound

This paper cites RepCodec: A Speech Representation Codec for Speech Tokenization.

WavChat: A Survey of Spoken Dialogue Models RepCodec: A Speech Representation Codec for Speech Tokenization

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.415495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.415495Z digest=sha256:e50d5c64d2c9a686fbb7accbcb5660d0b6187417c9c3cb69313bcb74a568c279

Observation 4a72e59f-fcf6-4b85-be7d-a80f72e08bc0 · outbound

This paper cites Residual Quantization with Implicit Neural Codebooks.

WavChat: A Survey of Spoken Dialogue Models Residual Quantization with Implicit Neural Codebooks

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.421599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.421599Z digest=sha256:f9193fdbb46c6a8ff43574c1634778af52e6dbdb72a291039a86bfabf3c85a72

Observation 297bc52d-5e5f-46a9-9a58-0ef872718c9a · outbound

This paper cites The lj speech dataset.

WavChat: A Survey of Spoken Dialogue Models The lj speech dataset

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.427668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.427668Z digest=sha256:d94e5863edd5599ecd0086f2b803963eef643c6586b824c11165155376fb4588

Observation 22c33a22-114c-4876-b8a3-46a984a0ee61 · outbound

This paper cites Language-Codec: Bridging Discrete Codec Representations and Speech Language Models.

WavChat: A Survey of Spoken Dialogue Models Language-Codec: Bridging Discrete Codec Representations and Speech Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.433148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.433148Z digest=sha256:73c129b0007ef241173942f9961bb0a5ba8c30f1d9a7349ae8d7fc503c9cdb86

Observation 606ba9b5-f3b4-4a90-bb14-e3bb82890f64 · outbound

This paper cites MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech.

WavChat: A Survey of Spoken Dialogue Models MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.439022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.439022Z digest=sha256:478abf071c301cb5641975febacfb2015a8e8457ce0724ec317651e39f011d0b

Observation bd7dced9-e31f-430f-9d86-819999d292a9 · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

WavChat: A Survey of Spoken Dialogue Models WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.445227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.445227Z digest=sha256:e4c9e90d3347c2b021bfbf48a93d6e97e65fb08d66c766c77595b52056662ba0

Observation 162d83f8-3c4e-496a-9284-758657d4296e · outbound

This paper cites Textrolspeech: A text style control speech corpus with codec language text-to-speech models.

WavChat: A Survey of Spoken Dialogue Models Textrolspeech: A text style control speech corpus with codec language text-to-speech models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.452414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.452414Z digest=sha256:fcd3eaa2361a180a1ee00e2ef7be41018ade154296eb057b5194a10c9673ac97

Observation cf7bd55b-9218-4dcc-9454-8da0c777e3c1 · outbound

This paper cites ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control.

WavChat: A Survey of Spoken Dialogue Models ControlSpeech: Towards Simultaneous and Independent Zero-shot Speaker Cloning and Zero-shot Language Style Control

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.458611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.458611Z digest=sha256:a07d12a228f743b192fee09c74c8232c29cd9ed0eccf8cfa91bafa854df12cf4

Observation 4339d445-0079-4168-abe7-9884393aefd6 · outbound

This paper cites CVSS Corpus and Massively Multilingual Speech-to-Speech Translation.

WavChat: A Survey of Spoken Dialogue Models CVSS Corpus and Massively Multilingual Speech-to-Speech Translation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.464641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.464641Z digest=sha256:661d3a86c14dc84f65a0ca027e4f8ee37dca6ce7b496523a07c2c2f08bae509b

Observation 6cbdc066-4c28-4f04-949c-950e41ec9983 · outbound

This paper cites Mixtral of Experts.

WavChat: A Survey of Spoken Dialogue Models Mixtral of Experts

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.471085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.471085Z digest=sha256:7fc611dd5f666ce105202d1d5bc03e6edd98c7667c6290e795f0a823275ad882

Observation d82b99ff-d4cf-4386-8c52-5efd5a7850d2 · outbound

This paper cites Mega-tts 2: Boosting prompting mechanisms for zero-shot speech synthesis.

WavChat: A Survey of Spoken Dialogue Models Mega-tts 2: Boosting prompting mechanisms for zero-shot speech synthesis

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.477265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.477265Z digest=sha256:cfba3b95819be46b1ddf6d16434cf1f2101a0917b709db1abfb07a8666226619

Observation f123214f-8997-456f-9770-df9e795a12df · outbound

This paper cites Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias.

WavChat: A Survey of Spoken Dialogue Models Mega-TTS: Zero-Shot Text-to-Speech at Scale with Intrinsic Inductive Bias

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.482560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.482560Z digest=sha256:0c3f28fe201ff0a479611d0b5dda15f996f93d8cf1506ad26b1baa1e753442e7

Observation b63a201b-1a26-444b-89ed-9f5f2bd1ded3 · outbound

This paper cites Duplex conversation in outbound agent system.

WavChat: A Survey of Spoken Dialogue Models Duplex conversation in outbound agent system

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.488402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.488402Z digest=sha256:a11114925c56b9ee2eafdb5ba8d54d6ff436a658ffa57bc902c2cf60a326d51d

Observation fad135ec-e442-419b-929a-63b5ea94ab09 · outbound

This paper cites Efficient multimodal large language models: A survey.

WavChat: A Survey of Spoken Dialogue Models Efficient multimodal large language models: A survey

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.494016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.494016Z digest=sha256:98e1f270457342da87b2098be629cd9f8fb71db6733affe1495f3ffd51fa4a41

Pith citing papers

Observation a03dbb64-8ad8-47ff-8b52-f7946be435e7 · inbound

Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners cites this paper.

Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners WavChat: A Survey of Spoken Dialogue Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T21:12:24.013748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:12:24.013748Z digest=sha256:d08fcdff2f88a608d23f905e6a7997440834f0cb65ae6ee6e5b8bb70e9f2f4e6

Observation 85f77c63-dc55-4ddb-942b-4ab0570238d1 · inbound

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training cites this paper.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training WavChat: A Survey of Spoken Dialogue Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.935097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.935097Z digest=sha256:72a06b54e0ee89f3a15c8840f7c019e4be24d4892381a58dbfc53a430fffd799

Observation c9d119e6-23bf-4caf-a8a9-8ff3c6c41e7c · inbound

Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models cites this paper.

Advancing Singlish Understanding: Bridging the Gap with Datasets and Multimodal Models WavChat: A Survey of Spoken Dialogue Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T22:40:25.442576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:40:25.442576Z digest=sha256:e45fe51235f04f68128de7d003ceebd096784ede5d957aec1b96d67ded15bd39

Observation ac489d20-9269-4d47-9b1f-015ee9aef61b · inbound

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios cites this paper.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios WavChat: A Survey of Spoken Dialogue Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.952630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.952630Z digest=sha256:f909009844ed64ad861a35ae5eab55488bd3b21fc2e23a5e1940a7a10d3f8a4f

Observation 8cae7078-9d07-49f5-b004-9cd8c21c982b · inbound

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction cites this paper.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction WavChat: A Survey of Spoken Dialogue Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.898917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.898917Z digest=sha256:8752c638dce2f590df4387dde9679b0727208a17711f78fa305db7c76c69a239

Observation a5207cc1-6394-4a24-acb1-e3f238b866c6 · inbound

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey cites this paper.

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey WavChat: A Survey of Spoken Dialogue Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:36:19.156715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:36:19.156715Z digest=sha256:9cd3dcf2d48ddf3d02e5c44b5b3e8adfe4dd3bd3373bae29e8fadf6cddf96539

Observation 30ec01bb-3d89-4e32-b02f-f0bc7006d77e · inbound

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her cites this paper.

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her WavChat: A Survey of Spoken Dialogue Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T13:34:25.670790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:34:25.670790Z digest=sha256:cbf35865d969b0fd15a581d6876e3f8994dc4cac33b5a12666aac061015f0740

Observation d76f053a-4f51-44c6-9c0d-6b8f7b739e43 · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey WavChat: A Survey of Spoken Dialogue Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:45:08.048139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:94f46cb5e381b2faf3e2fce5cfea0fd0fa83396c4d14c9a6f1dfd31a5c345ffa

Observation ffc30f93-d563-4ea0-9351-857f2670ad54 · inbound

LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis cites this paper.

LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis WavChat: A Survey of Spoken Dialogue Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:52:07.865325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:52:07.865325Z digest=sha256:69b077ef271be2d515f8f37b1131ca54bcb690eb09087a19a700cec7b043ee16

Observation fd7ba0be-36c5-4455-8098-6c4bf2e123b0 · inbound

PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs cites this paper.

PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs WavChat: A Survey of Spoken Dialogue Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:00.748722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:00.748722Z digest=sha256:25928d03c08e7cc01e8c45a8de782ab7326f183528125eb7c9f6d555f01a5e61

Observation 24297cbb-faa3-4c2a-a25a-9cf84fe8d9b4 · inbound

Speechless: Speech Instruction Training Without Speech for Low Resource Languages cites this paper.

Speechless: Speech Instruction Training Without Speech for Low Resource Languages WavChat: A Survey of Spoken Dialogue Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:15.917688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:15.917688Z digest=sha256:bdd274637ae3f688c5c0519859eb50247f9d05331dfeefebfd11370e9ea73b8b

Observation 4c3dccc7-2e75-4285-ad77-48527129a280 · inbound

BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM cites this paper.

BR-ASR: Efficient and Scalable Bias Retrieval Framework for Contextual Biasing ASR in Speech LLM WavChat: A Survey of Spoken Dialogue Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:09.865471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:09.865471Z digest=sha256:97dbbf6bccf3ee55054537e0744866bb0ad839bbef2895b7a621d863e4916b30

Observation b3c742b7-aa8e-4ea6-b906-76cf35b79278 · inbound

Breaking the Barriers of Text-Hungry and Audio-Deficient AI cites this paper.

Breaking the Barriers of Text-Hungry and Audio-Deficient AI WavChat: A Survey of Spoken Dialogue Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:51.286122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:51.286122Z digest=sha256:b95b412aa7efff9105f58d01c1d4b9c085290a744ed58a897e7423eb2d3d9e0a

Observation 44080555-c607-49e4-a760-a70df34c9de0 · inbound

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training cites this paper.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training WavChat: A Survey of Spoken Dialogue Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:11.320572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:11.320572Z digest=sha256:c90da84810b95ff502fb131c7eb6fba1ba2457efd2833fcfeb57b27dfd847d49

Observation 681a5cdb-7ba7-4382-9e1f-b936d7d04130 · inbound

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model cites this paper.

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model WavChat: A Survey of Spoken Dialogue Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:13.393147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:13.393147Z digest=sha256:d4d1a03efcca8a21ef1c273c7d36f3f4d93873d8d91a1544761fc683463f85a7

Observation b65c5f37-2a5f-4cad-b67c-2cbb5e689769 · inbound

OpusLM: A Family of Open Unified Speech Language Models cites this paper.

OpusLM: A Family of Open Unified Speech Language Models WavChat: A Survey of Spoken Dialogue Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:36.084747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:36.084747Z digest=sha256:7dd443a22bc5e1a2ee4a729ab6a3b89d79590110ba5e9319e09d26c9b2190d13

Observation ae0594c1-2ab7-40ed-84a5-211ae59b4860 · inbound

UniMind: Unleashing the Power of LLMs for Unified Multi-Task Brain Decoding cites this paper.

UniMind: Unleashing the Power of LLMs for Unified Multi-Task Brain Decoding WavChat: A Survey of Spoken Dialogue Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:47:10.364647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-19T07:43:26.736409Z digest=sha256:8d1afa3929c93070be49ae8668ba230de87e1538eda7330916e66b76caf70c9e

Observation 8c503e55-8ba0-4943-9f41-c006f5503240 · inbound

StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding cites this paper.

StreamFlow: Streaming Flow Matching with Block-wise Guided Attention Mask for Speech Token Decoding WavChat: A Survey of Spoken Dialogue Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:33:09.245104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:33:09.245104Z digest=sha256:42b3ae14c5c245b584391fe0d5f34b1d12f381ea066f7a2fd33eb56b423b4742

Observation f140ba98-4637-4a27-b434-b70b112dc254 · inbound

Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling cites this paper.

Towards Accurate Phonetic Error Detection Through Phoneme Similarity Modeling WavChat: A Survey of Spoken Dialogue Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:11:24.285128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:11:24.285128Z digest=sha256:580a61f9aca04f4c2493413e0ec6f595c4621d0625afde6e9d78dbd34a93445f

Observation 74403b0a-948f-405d-9850-ed7df164adbb · inbound

Dual Information Speech Language Models for Emotional Conversations cites this paper.

Dual Information Speech Language Models for Emotional Conversations WavChat: A Survey of Spoken Dialogue Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T21:45:22.176873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:45:22.176873Z digest=sha256:d83400eb31a47089acc3de3755f67f5d427037fda1b6de70e1449561753b75b9

Observation 5288e3c4-2220-4f74-bd8a-ccf2e787a69d · inbound

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs cites this paper.

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs WavChat: A Survey of Spoken Dialogue Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T16:46:49.184979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:46:49.184979Z digest=sha256:e94e1382483c66f5a488f98c9952c311272075c95e01909f4ca790bfb653e550

Observation 9f744dbb-9fdf-44a8-9b65-209172234df5 · inbound

ChipChat: Low-Latency Cascaded Conversational Agent in MLX cites this paper.

ChipChat: Low-Latency Cascaded Conversational Agent in MLX WavChat: A Survey of Spoken Dialogue Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:44.007778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:44.007778Z digest=sha256:55c217ff851014e39a090ad394969a279e7c541eb7845e2586d3e5125dc0883d

Observation ab9084d7-ed4c-4928-add3-05e6c2c8fd8c · inbound

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models cites this paper.

Game-Time: Evaluating Temporal Dynamics in Spoken Language Models WavChat: A Survey of Spoken Dialogue Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:52:35.726992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T11:51:43.561210Z digest=sha256:b72e3a073abdb8b561fd0b589f622875e4e12e915204dac9a221464aa58cacc5

Observation b19cbf1b-6668-4905-b09c-0e80e2297a68 · inbound

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM cites this paper.

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM WavChat: A Survey of Spoken Dialogue Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T21:42:49.032029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:42:49.032029Z digest=sha256:9fae4e592723ca1ea19030e004afedca6c15c06f518895f50dca3da3a8618d9a

Observation 271fc78e-d208-4431-9e80-c67457bd918e · inbound

TiCo: Time-Controllable Spoken Dialogue Model cites this paper.

TiCo: Time-Controllable Spoken Dialogue Model WavChat: A Survey of Spoken Dialogue Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:39:35.866368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T00:38:52.182973Z digest=sha256:227c58115be86c5aa5960925b7b988042a297f02cd1217a7e9f5b790c2abfc46

Observation 535b2d08-65a5-4aa4-ba7a-7e720b089a69 · inbound

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff cites this paper.

A Synonymous Variational Perspective on the Rate-Distortion-Perception Tradeoff WavChat: A Survey of Spoken Dialogue Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T20:09:57.922788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:09:57.922788Z digest=sha256:e22d0ae4acf83800cf4b146d9fa6c53321eb3004b8e6d4c71c4193aaa75b91ae

Observation 7bd73e9c-b565-485d-acf0-d87c0ab92ffc · inbound

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection cites this paper.

Hijacking Large Audio-Language Models via Context-Agnostic and Imperceptible Auditory Prompt Injection WavChat: A Survey of Spoken Dialogue Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:35:18.904751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T11:32:10.126062Z digest=sha256:c60e5dcd9355046a595f19a3e93c8732f2dc0d86d3e00b9d381c66d2a7063868

Observation 8ea7f99f-86af-441f-a610-0d9a707b88a3 · inbound

Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals unreliable Multi-Turn Behavior in LLMs cites this paper.

Talking to a Know-It-All GPT or a Second-Guesser Claude? How Repair reveals unreliable Multi-Turn Behavior in LLMs WavChat: A Survey of Spoken Dialogue Models

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T02:22:21.217967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T02:18:07.162514Z digest=sha256:ad7c4549b469109210dfb4a2af409396083485512427bef54499f197c4f2a6ad

Observation e8014a72-e25e-48c7-9938-bfde1e7c5e27 · inbound

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models cites this paper.

Walking Through Uncertainty: An Empirical Study of Uncertainty Estimation for Audio-Aware Large Language Models WavChat: A Survey of Spoken Dialogue Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:41:26.363121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-07T14:29:18.348031Z digest=sha256:b879ddd164791c08938b183dacbecce655ddca8aeefd4f7c8b679ac704b51ecb

Observation 6de18e82-e85e-4747-9da7-4f588149be8f · inbound

VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models cites this paper.

VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models WavChat: A Survey of Spoken Dialogue Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:51:09.764621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T16:59:25.973854Z digest=sha256:9ba939001c59107a9457f5d46e04879b10e90ffcdb102474878b92a6eaabeb39

Observation 36ee66e8-0806-429d-94e4-0367b3d78a91 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing WavChat: A Survey of Spoken Dialogue Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:50:55.760212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:80de40dda671ffac2a71dcf112b97e874a8eda36a31c789103ddff10d9399198

Observation 7a2fed10-bd5c-417e-bed0-684bfb935262 · inbound

Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades cites this paper.

Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades WavChat: A Survey of Spoken Dialogue Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:03:17.758837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T13:02:53.960375Z digest=sha256:240a89f19e9f2b383f59537f01690c1671ccdccb13d4c420d6db58cf6782c347

Observation 6d257d6d-1135-4f57-aec4-5ffc20237160 · inbound

Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades cites this paper.

Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades WavChat: A Survey of Spoken Dialogue Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:15:00.392581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T19:12:32.672352Z digest=sha256:18c0fd774eb32e9788e8ff8259f1a29b998de50c4f68dc0e314c17f572d138b1

Observation 3e1a5b93-42aa-42d3-b65a-0f686d7bb436 · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook WavChat: A Survey of Spoken Dialogue Models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:49.050610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:43154630fc9ad5b36f8e0bec698f6026456d6c58f7de97a7b6a37fdba677f5b3

Observation 292c42a6-2bcc-4fdc-b4d3-5472e442d172 · inbound

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects cites this paper.

PolySpeech-100: A Large-Scale Benchmark for Speech Understanding Across 100+ Languages and Dialects WavChat: A Survey of Spoken Dialogue Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:13.838719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T17:44:07.669223Z digest=sha256:a0dc3bd6ac4cff5adaa9a3fce79314ef9b9f965b39eb3c5789778c03f4da3e34

Observation 56dd88e1-fd28-4878-aafe-63dad0995384 · inbound

UniVocal: Unified Speech-Singing Code-Switching Synthesis cites this paper.

UniVocal: Unified Speech-Singing Code-Switching Synthesis WavChat: A Survey of Spoken Dialogue Models

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-07-02T00:46:24.646473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T13:17:13.510587Z digest=sha256:85922a6cf549d773d41d9c10a30df9c585cfd39968524a814b8a509c56b7dfde

Observation 9e561681-2cfb-4359-be47-55f7f9d8f1e8 · inbound

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models cites this paper.

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models WavChat: A Survey of Spoken Dialogue Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:20:57.380720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T12:53:00.569655Z digest=sha256:79d9b88f32252d000568251fd7397d67ae8d7d057cb2b7e9a7078c462ce560fd

Observation 5615766a-21f1-4d73-8a28-d4a93a13076e · inbound

Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering cites this paper.

Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering WavChat: A Survey of Spoken Dialogue Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:39.504681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T13:22:52.428152Z digest=sha256:c7c85338e4292d767fb987834678544664752c268642d29b69a48f1d5f2c139d

Observation fc803355-0644-4879-9499-af55bdc620cc · inbound

PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue cites this paper.

PRISM: Prosody-Integrated Multi-Agent Reasoning Framework for Empathetic Spoken Dialogue WavChat: A Survey of Spoken Dialogue Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:33.547578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T06:53:13.344795Z digest=sha256:bdafdba4b52c8e5d2ac93c79613cf6271f4d9384edd2f95d89da9e86173afa23

Observation 0a6e3651-30ce-49cd-a0a5-a1d5c29ac73a · inbound

Integrating Facial Generation into Full-Duplex Spoken Dialogue Systems cites this paper.

Integrating Facial Generation into Full-Duplex Spoken Dialogue Systems WavChat: A Survey of Spoken Dialogue Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:41.319388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T12:08:55.070482Z digest=sha256:dbc18a37f59b65b0a0f24815c81447e46db389e6f1ebf1b26e5a3c544e8297e7

Observation 5bff76c7-2b3a-478d-bc3e-aaf2a390e528 · inbound

Audio Editing in the Era of Foundation Models: A Survey cites this paper.

Audio Editing in the Era of Foundation Models: A Survey WavChat: A Survey of Spoken Dialogue Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:09:49.441309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T07:10:49.047784Z digest=sha256:e11df01cc9e817541e33f3e19b17bc65cf66512e28feb77b34cd0c3489b2458c

Observation c068f586-cd16-4bd7-9cdb-a91e167203ce · inbound

RedVox: Safety and Fairness Gaps in Speech Models Across Languages cites this paper.

RedVox: Safety and Fairness Gaps in Speech Models Across Languages WavChat: A Survey of Spoken Dialogue Models

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:52.893021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-26T04:37:00.399470Z digest=sha256:78e8396622a2e4b61240fe0873e5023a9756590895161c82de76b1ec90487703

Observation 16ed510f-853e-45a6-b971-8f702480c31a · inbound

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue cites this paper.

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue WavChat: A Survey of Spoken Dialogue Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:28:58.297693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-03T21:22:38.888528Z digest=sha256:f2026dffd70f5dc5b16bcae0aa30aba34a85353e20ccf0bc35237ae6ee59cb3f

Observation 1aacc01f-e37b-4fce-9c64-46d6f0f3761f · inbound

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue cites this paper.

TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue WavChat: A Survey of Spoken Dialogue Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T08:53:13.408944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:53:13.408944Z digest=sha256:be3a98331fcb892f007958c2c2df7827b11ee1a418d00de7d648ead1d82d1d1d

Observation b65231ce-691f-49ab-bddc-1b16309664a0 · inbound

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs cites this paper.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs WavChat: A Survey of Spoken Dialogue Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.611257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:3b6bb3de98fe5415314898e867764bb669a0f74c287d46fa72389d42d685051d

Observation 26ce0f7a-9521-40c9-a8ca-99e9b2e5484c · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment WavChat: A Survey of Spoken Dialogue Models

Reference 131

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:17.646123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:17.646123Z digest=sha256:153bc71b4bff8aec4d05af8d64a5db4a4da696428c8ed210acfe2966be9819db

Observation 773619b2-f9bc-4a5f-aa97-df3bf5dc5ab0 · inbound

VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference cites this paper.

VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference WavChat: A Survey of Spoken Dialogue Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:59.509929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:59.509929Z digest=sha256:8ba1e1c43eaafadeb7837803a891b287b8f81a12746283970c55600cf5effd1d