Pith. sign in

Paper Citation Record · LEDGER

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems

As of 21 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 2 inbound Pith citation observations for arXiv:2506.00722.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00722 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:04:33.908292Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:04:28.710605Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:04:34.616301Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy37
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3275385d-1d06-4ed2-8613-0cdec743b139 · outbound

This paper cites De- spite their growing importance, building effective spoken dia- logue systems remains a challenging task due to the complexity of human communication.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems De- spite their growing importance, building effective spoken dia- logue systems remains a challenging task due to the complexity of human communication

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:39.408493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:28.638049Z digest=sha256:18e51c1b42fcb428381533cd285953f18d37dd3f422d4dfedc186d8ac85ffcbb

Observation 5fffcd6a-6b6d-4969-a009-cc3ca9c495d5 · outbound

This paper cites Chain-of-Thought Training for Open E2E Spoken Dialogue Systems.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Chain-of-Thought Training for Open E2E Spoken Dialogue Systems

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:04:34.675702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:28.710605Z digest=sha256:171394942dd25984516308260c9a43a518ed1eab2992df2d8b6f3983654cfc6d

Observation af1bdbe0-775b-45a9-9b10-a04dfaab3ac2 · outbound

This paper cites tar- get.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems tar- get

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:04:39.324937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:28.818087Z digest=sha256:24a9eef2358f3d1ff12b04b90a5252f25408ab1d46a1a64a0b5aef5a59565985

Observation 7f383818-59a6-4a87-9a7e-97b0aa8c134a · outbound

This paper cites Spk prompt.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Spk prompt

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:39.221087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:28.894332Z digest=sha256:fe46f43502f219dfd2749a75275ef4537331e9719dd4c3f792c37ae05bb4661d

Observation afa856c3-a135-46ca-929e-57e1462259e6 · outbound

This paper cites SpeechLM E2E.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems SpeechLM E2E

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:39.117633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:28.982116Z digest=sha256:00e18b76db6ee6f25122be5d3eae24fce11e0fc6ca532860c873d1f47d15817c

Observation 70b3e8e3-7e1a-49c7-9d27-3dabb78b0cf4 · outbound

This paper cites speaking while listening.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems speaking while listening

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:39.025377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:29.048752Z digest=sha256:e20caaf6b5f40e1d80e9f76eb3b518ae2e3a9e78bfc16bb8a4f187dbb0aa2210

Observation 8d259898-2dec-4611-a3ce-c87381376359 · outbound

This paper cites an unresolved cited work.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:04:38.929616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:29.145753Z digest=sha256:8e464cb93c5a78d215e290a902d4f2e6a61e35e416af4fb2f51386532847f751

Observation b2c97fe5-98a3-4316-9d69-d2c50c8ea88d · outbound

This paper cites Jokinen et al.,Spoken dialogue systems.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Jokinen et al.,Spoken dialogue systems

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:38.869816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:29.250887Z digest=sha256:2b4ee72d86e063485528395888a817b7b236f689ae9ecb9089ac9bcf81b29e52

Observation a079a368-b231-47f0-91b7-c43cdf491c52 · outbound

This paper cites Social robots that interact with people,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Social robots that interact with people,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:38.776514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:29.396887Z digest=sha256:c2bf63c5a7be043742124b509d88074d370623e7f70bbb1fe4b6557ea5b9e0b3

Observation 8472b116-2319-4b95-87d9-a369c48c10a4 · outbound

This paper cites Challenges for spoken dialogue systems,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Challenges for spoken dialogue systems,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:38.692861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:29.504095Z digest=sha256:c6236d761d3a74e94e348478aa60343e7b9e921d6d6f4c562bc3fc8944d47027

Observation 514ec6ee-90e9-4dfe-8809-956b2a07f4f0 · outbound

This paper cites Audiogpt: Understanding and generating speech, music, sound, and talking head,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Audiogpt: Understanding and generating speech, music, sound, and talking head,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:38.597766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:29.586574Z digest=sha256:75f4ed575dc1fb9d1b570791272a6b445af3c6e133e4330c73384cf0ea81666b

Observation 3b013b1b-f082-4a50-99fd-6b676cd1ea02 · outbound

This paper cites Wiseman,Py-webrtcvad, Accessed: 2024-12-10, 2024.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Wiseman,Py-webrtcvad, Accessed: 2024-12-10, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:38.519634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:29.669640Z digest=sha256:10c6bf3eebbc20b94e6763b965be1eded2ee0cf49f2e0dcde7d4c679e1488555

Observation 06609268-711b-450f-8a9a-47370e6c77d8 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Robust speech recognition via large-scale weak supervision,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:38.429081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:29.751567Z digest=sha256:f7da57bec24d036a57957f3b80668ac6611103e3549db32e4bdafebca678642c

Observation 7f481491-5ca4-4e27-990f-76770b03126c · outbound

This paper cites Reproducing whisper-style training using an open-source toolkit and publicly available data,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Reproducing whisper-style training using an open-source toolkit and publicly available data,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:38.362200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:29.822980Z digest=sha256:c88211eee7c907ba757f1d3c9d87030379b990de8a9f3ea28f00dbb3f25e7c73

Observation 3ffa3e26-6c9e-48d2-980b-135322202982 · outbound

This paper cites DialoGLUE: A Natural Language Understanding Benchmark for Task-Oriented Dialogue.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems DialoGLUE: A Natural Language Understanding Benchmark for Task-Oriented Dialogue

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:29.893689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:29.893689Z digest=sha256:f5f2abed05d37d4b21e65a95781fe200f3dece93f969454409bf5fd50b1b5ec8

Observation f22beea7-eb57-4278-b3f3-8187c0f29798 · outbound

This paper cites Few-shot natural language generation for task- oriented dialog,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Few-shot natural language generation for task- oriented dialog,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:38.297387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:29.989176Z digest=sha256:8968d98e80c4f323f1cf4b53aeac7f927a9cf879b4cbfb58e1457837ba3ce13e

Observation 180a7f64-491d-4f52-ad85-7dae5c7f94df · outbound

This paper cites an unresolved cited work.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:04:38.229983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:30.088918Z digest=sha256:63ab3700a6d04a3820c273529f3830c67569f79d53b45432cf5588bba83a0330

Observation 2f8f7c4a-e6c6-4d0f-b914-11d8bf95a6c8 · outbound

This paper cites Switchboard: Telephone speech corpus for research and development,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Switchboard: Telephone speech corpus for research and development,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:38.127983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:30.193100Z digest=sha256:2d8ec7d909fb2257b9a1d98b47d84dcc483a72c33b5268f484925f3c9bea1a3c

Observation 95fc4281-880e-4086-8e0f-1fa147969e57 · outbound

This paper cites Towards Empathetic Open-domain Conversation Models: a New Benchmark and Dataset.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Towards Empathetic Open-domain Conversation Models: a New Benchmark and Dataset

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:30.292162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:30.292162Z digest=sha256:cb20118dca50176ea37c0f3a1a5060f5f6cb9ee305b20184e7096e21c6703ddd

Observation 0fecf16e-49b4-4a13-a455-e77b8ff2b8c8 · outbound

This paper cites Prediction of turn-taking using multitask learn- ing with prediction of backchannels and fillers,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Prediction of turn-taking using multitask learn- ing with prediction of backchannels and fillers,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:38.058133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:30.378219Z digest=sha256:ebd8fc142593a7cd8a631840703b1bb34c4c9001d1940218590e81878db4cc9b

Observation 582cf38a-c8f3-45d9-92f7-7d0a489993ae · outbound

This paper cites Prosodic features which cue back-channel re- sponses in english and japanese,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Prosodic features which cue back-channel re- sponses in english and japanese,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:37.966742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:30.483496Z digest=sha256:fb0d5f7bd0b663371dd6198897539909964189c8180b871e0a7aac6ebdcf8b53

Observation 47ca166e-e7cd-480c-8db5-630e5d806b0a · outbound

This paper cites Automatic acoustic synthesis of human- like laughter,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Automatic acoustic synthesis of human- like laughter,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:37.907951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:30.568396Z digest=sha256:5bf4efc3230647ebcc8e9ca12790b63b68bf2bd1079f26b4527d27fb3cd3f6d9

Observation 9962d161-8fb8-4812-90ed-a8086abefb5d · outbound

This paper cites A conversation robot with back-channel feed- back function based on linguistic and nonlinguistic informa- tion,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems A conversation robot with back-channel feed- back function based on linguistic and nonlinguistic informa- tion,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:37.849661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:30.633919Z digest=sha256:e85c56afd4f261375cedc0edbaa3bd7339f3d9a2a7b8cc25c50bed8bca8b2201

Observation fb68686b-f3e2-40f9-a196-7ecebfb1fc38 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:30.764440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:30.764440Z digest=sha256:38e19b6e8e0766f24d8c42bc598b8e64f22d506de212e2e88bf5b63abd49de0b

Observation 67f1c7fa-4684-4f89-8cd1-5f1bef68f736 · outbound

This paper cites Xie et al.,Mini-omni: Language models can hear, talk while thinking in streaming, 2024.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Xie et al.,Mini-omni: Language models can hear, talk while thinking in streaming, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:37.808636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:30.863948Z digest=sha256:3150f6f12730a409b3477b013bd7674af696ad2f6ca8474d2543e209e93ba318

Observation f772bbac-eeff-4ee5-9d4e-38c156476194 · outbound

This paper cites Moshi: A speech-text foundation model for real-time dialogue,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Moshi: A speech-text foundation model for real-time dialogue,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:37.750467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:30.967213Z digest=sha256:302825c408cfa6d9db4fa164d3f71b6535b3b81918e92f3eddea79412c34c9dd

Observation eaffa489-33da-48a6-b95d-88405054ab5a · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Chain-of-thought prompting elicits reasoning in large language models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:37.719663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:31.052552Z digest=sha256:0c51dc7c4d4d889b73bfe63a4812db86ca6f5a1fc5440dac872a54f688309846

Observation b0695da3-f98b-49b1-8eb9-cef407085ef6 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Automatic Chain of Thought Prompting in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:31.120330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:31.120330Z digest=sha256:3fede2a8b48154408d9a694d917a75543ae11c60528f55a0c287c8f41a194edf

Observation a1903272-164d-427e-bbcb-05979ce78046 · outbound

This paper cites ESPnet-SpeechLM: An open speech language model toolkit,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems ESPnet-SpeechLM: An open speech language model toolkit,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:37.688594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:31.235594Z digest=sha256:8e8abfe4219a0c4a9f7401b28d1eb4b012c47ba2d54b2b95ab43b7876efd4c02

Observation 75b37a5b-155b-4e53-8385-de08dc27b3bf · outbound

This paper cites Ji et al.,Wavchat: A survey of spoken dialogue models, 2024.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Ji et al.,Wavchat: A survey of spoken dialogue models, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:37.654547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:31.334137Z digest=sha256:d689c1d98e2db2d5e92ebfd4bd734572114016dc1656bbad6aaa4b54e0242655

Observation eb42478d-6c08-4a06-b236-b8db484a2420 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:31.419862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:31.419862Z digest=sha256:fef4b413345bded8a38aa3549c3bd3e23c95dc743927d672ec584da36fe5ceea

Observation 8e9d583b-8ef8-4035-a643-3f27cd917c3a · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:31.519528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:31.519528Z digest=sha256:2f868e30f84cf64ac7a430d2dd88561513d7108cc9d1c8a9bfe9756daa84a715

Observation e9c68199-96f5-4c3a-8ce6-0ca1a4b2e412 · outbound

This paper cites Distilling an End-to-End Voice Assistant Without Instruction Training Data.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Distilling an End-to-End Voice Assistant Without Instruction Training Data

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:31.596425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:31.596425Z digest=sha256:2284d81b4f334592952709820eb468d720b18403951f6cd755356eb8fdedbe15

Observation f639d97f-b435-4472-8113-e18b01eb2ec4 · outbound

This paper cites SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems SpeechGPT-Gen: Scaling Chain-of-Information Speech Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:31.706485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:31.706485Z digest=sha256:bbd886028e985bc22b725316d4636a072e8de02ae6f5d396e03a8269370dfa3f

Observation d8644870-f3ed-4279-bc91-0e79b8f81c5c · outbound

This paper cites Generative spoken dialogue language mod- eling,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Generative spoken dialogue language mod- eling,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:37.633454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:31.796410Z digest=sha256:9a22586f93ef8afc3097e34a5424cf87b995e106b81ae97412d4ea7db93fdbe9

Observation 60021f98-1691-4f71-81f2-fde300de9fd8 · outbound

This paper cites Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Beyond Turn-Based Interfaces: Synchronous LLMs as Full-Duplex Dialogue Agents

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:31.862397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:31.862397Z digest=sha256:02ebe792a7a7cee3b53c25263970c10f890d12b90119035c1322ef590ad31eff

Observation ce5592c6-903f-44ff-ae2d-861a45d3c6c1 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:31.966356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:31.966356Z digest=sha256:14396dbcc514ae44de6eb0e941b847bb5f8dbd985cf360c14144345ee333aca4

Observation 214678fc-2ea0-4cd8-a8c5-dd8f580ac8c3 · outbound

This paper cites Parrot: Autoregressive spoken dialogue lan- guage modeling with decoder-only transformers,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Parrot: Autoregressive spoken dialogue lan- guage modeling with decoder-only transformers,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:37.588627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:32.062938Z digest=sha256:260fccd03deac4318615966d67ceb886b7c53cce16742b10ce708de7e058996a

Observation eb7c36ed-3245-4565-994f-66586859dc67 · outbound

This paper cites OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:32.164735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:32.164735Z digest=sha256:9958c751cf9300a4f3cad8930465a096f99ff7450a3912caaabaaf3ef0dd84c5

Observation 47ee95e5-9b67-404d-876b-63204613dbb3 · outbound

This paper cites Audiolm: A language modeling approach to audio generation,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Audiolm: A language modeling approach to audio generation,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:37.386081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:32.266205Z digest=sha256:0dcda40e7c994aedf4dd72395ea6898c7b697dcdaf153f78bf06b49c62ff9bc3

Observation 2222b6f6-6cbf-4cef-b04f-68ad6c0c3849 · outbound

This paper cites The fisher corpus: A resource for the next gener- ations of speech-to-text.,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems The fisher corpus: A resource for the next gener- ations of speech-to-text.,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:36.985959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:32.358762Z digest=sha256:9a5c9f2128a76d995a5e64b7ada435ae24e8db7143e84735120282b76b090d91

Observation d97f2445-8715-493a-8ee8-d148d4e8905c · outbound

This paper cites Rouge: A package for automatic evaluation of sum- maries,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Rouge: A package for automatic evaluation of sum- maries,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:36.654739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:32.465673Z digest=sha256:cace1c1c2064daa750054bdd046baafea674d58575ecf92303a071fda2590c11

Observation cac99c06-27e6-4692-9a43-a2b6a703725b · outbound

This paper cites Meteor: An automatic metric for mt evalua- tion with improved correlation with human judgments,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Meteor: An automatic metric for mt evalua- tion with improved correlation with human judgments,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:36.373483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:32.665718Z digest=sha256:735542b5f76f7e129d83feafe4c890d9f82be86667df49e0dec0430d3e475de3

Observation 2eaad0fa-a80c-4e64-bba2-1534c9972aa7 · outbound

This paper cites Perplexity—a measure of the difficulty of speech recognition tasks,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Perplexity—a measure of the difficulty of speech recognition tasks,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:36.149602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:32.778357Z digest=sha256:e7dc47afb42d9d0862fb2726f9883ed9025df4833ce3223412ab2cfdefa99867

Observation 9c5b451e-04dc-41a6-999c-8e3bd1f95894 · outbound

This paper cites Language models are unsupervised multitask learners,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Language models are unsupervised multitask learners,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:36.031174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:32.845649Z digest=sha256:8921cb6c819dd95a085d0e48ca326165a8cae703b7b19802d449ae63080b5674

Observation bce3c563-952d-44de-802d-51e9c9e8c4c1 · outbound

This paper cites VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:32.968338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:32.968338Z digest=sha256:b57a90012b5bc254faa5357166424ff524bbd29f549d4fa2897f40369a9a7093

Observation 82344473-cb3b-46fa-8cf3-083ae7ee41df · outbound

This paper cites UTMOS: UTokyo-SaruLab system for voice- MOS challenge 2022,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems UTMOS: UTokyo-SaruLab system for voice- MOS challenge 2022,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:35.907626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:33.040888Z digest=sha256:cd14fdb4c8407eb7ca2adaff642bb3214d914dbec3c6c299aa53b318d4bc7872

Observation 6450d69a-9098-42c2-b047-1a05570eae17 · outbound

This paper cites Emotion2vec: Self-supervised pre-training for speech emotion representation,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Emotion2vec: Self-supervised pre-training for speech emotion representation,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:35.766771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:33.146326Z digest=sha256:2436338f541ad22553519faa47b65d8a3527876f9dc1bc72ec8ac47346da8b06

Observation 14252891-b28a-48bc-ae12-c02a804efde9 · outbound

This paper cites an unresolved cited work.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:04:35.616128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:33.208286Z digest=sha256:58913b027baf2a67d25c3ea72b24e9aa2f5e99e342eab87a45a9bbb70a3bcb47

Observation e05b9044-e744-4503-ad34-fa414f63e486 · outbound

This paper cites Espnet-tts: Unified, reproducible, and in- tegratable open source end-to-end text-to-speech toolkit,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Espnet-tts: Unified, reproducible, and in- tegratable open source end-to-end text-to-speech toolkit,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:35.475910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:33.278872Z digest=sha256:de62ea747c7f7bc3096962f8ad904d0c22ad04dbcfa769e1612a418430deac06

Observation 22873f7f-566a-4fd0-89ef-69f90fd64827 · outbound

This paper cites ESPnet: End-to-end speech processing toolkit,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems ESPnet: End-to-end speech processing toolkit,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:35.328041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:33.392059Z digest=sha256:52c85611f75d71ba2bbe972d17405edefa2911be2d2517cfef3451580a48fcff

Observation 837a7668-1b24-4819-8fc7-6e16cedccd22 · outbound

This paper cites Espnet-slu: Advancing spoken language under- standing through espnet,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Espnet-slu: Advancing spoken language under- standing through espnet,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:35.183696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:33.509515Z digest=sha256:ca86bb91b3ff3902f231b8728a3149b7d7493662bdb1a2a928da7e79ec8a4ecf

Observation 4b3d05f1-1b0e-4db2-af8c-02e4934e1445 · outbound

This paper cites Simple and controllable music generation,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Simple and controllable music generation,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:35.032799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:33.641917Z digest=sha256:baf9741b922bf6e2e443517b3d1a82f8b251a8e824d350c74d187d5d411eddfa

Observation 8a5f30cd-9596-4953-bc6d-a8f17d175927 · outbound

This paper cites ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:33.706834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:33.706834Z digest=sha256:216aee2dfb84c776ebf5755b924217dfc221a86a662aedec480df78fbf05b411

Observation e7ef973d-14e2-43c9-964e-96bbab9d46dc · outbound

This paper cites Towards Robust Speech Representation Learning for Thousands of Languages.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Towards Robust Speech Representation Learning for Thousands of Languages

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:33.822916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:33.822916Z digest=sha256:54af35c6aa34a63f920ad4d57da19337a13cb9606669ebd66758beb6ac3f4766

Observation e4a9e794-3aeb-409d-95c4-bb5094a10da0 · outbound

This paper cites ESPnet-SDS: Unified toolkit and demo for spoken dialogue systems,.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems ESPnet-SDS: Unified toolkit and demo for spoken dialogue systems,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:04:34.881160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:33.908292Z digest=sha256:b9dd100397011b3f1cb11057c1d9ba3645ae3e1dbdd5138994c1794876aba6bf

Pith citing papers

Observation 5fffcd6a-6b6d-4969-a009-cc3ca9c495d5 · inbound

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems cites this paper.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Chain-of-Thought Training for Open E2E Spoken Dialogue Systems

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:04:34.675702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T12:04:28.710605Z digest=sha256:171394942dd25984516308260c9a43a518ed1eab2992df2d8b6f3983654cfc6d

Observation 73327330-5df7-4140-8d91-b147b4471cc2 · inbound

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models cites this paper.

Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models Chain-of-Thought Training for Open E2E Spoken Dialogue Systems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T11:20:17.708968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:20:17.708968Z digest=sha256:a5f00bca8a890f32cc15714603b886f8a9776eeaad999fbdca578f7b4e8d7f4d