Pith. sign in

Paper Citation Record · LEDGER

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

As of 11 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 5 inbound Pith citation observations for arXiv:2501.01384.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.01384 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:32:05.104984Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T23:17:03.456746Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:17:28.830734Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6f033cc8-adac-4012-aae0-4cf27578d44c · outbound

This paper cites SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.897175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.897175Z digest=sha256:780cb831e9ca1ed0e3ca54b0638edec319c16538c280b7426366b3615b2a718a

Observation 0d6982a2-999e-4f38-8807-0b918d657330 · outbound

This paper cites Qwen2-Audio Technical Report.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Qwen2-Audio Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.927050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.927050Z digest=sha256:29b4cdbff9445dd351a4ae19cf819fbfb51b8814f56fcc960d2c1805e55404bf

Observation ac489d20-9269-4d47-9b1f-015ee9aef61b · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios WavChat: A Survey of Spoken Dialogue Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.952630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.952630Z digest=sha256:9177a958d469dd42dd44d7c2332a552c72d42400361bf3f97f919e2f078fb4f1

Observation eea48f33-c7b0-4afa-98f3-7c95a493b137 · outbound

This paper cites Dailytalk: Spoken dialogue dataset for conversational text-to-speech.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Dailytalk: Spoken dialogue dataset for conversational text-to-speech

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.788900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:32:04.973121Z digest=sha256:256cee6ad6060affd4b4d6457634f65f40ce5cbe1e9ff143fe3b3ee92bf7d389

Observation a1bda65c-69da-42a3-9bee-ba9a7428ecda · outbound

This paper cites Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.980466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.980466Z digest=sha256:ca38c238d289518a6f09ee6302e878e376ce30002e975718fdbd6174518c0301

Observation cab6075e-e96e-4126-8053-974c7a93de7b · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.987198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.987198Z digest=sha256:d189c0a7044e2d50e4da161e46ea6b568f8d57c624995437c71eb6f0a1d11bd5

Observation ee9a3be0-a176-45b5-bda9-b2bdf0e8327e · outbound

This paper cites EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.997721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.997721Z digest=sha256:6dc4d9cf2f34b0bac615a290b92857f991b2b41c9c61afd76693aba5289ae9b0

Observation afe84b7f-9da2-47d1-bc17-7dabddfe33cb · outbound

This paper cites MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.020670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.020670Z digest=sha256:fa08ae7b0d70736216dbd8abc6c6567a424f1b40eb52c135e2c54d38fa32d9a0

Observation 9671ed34-5a79-4e7e-bbbb-c4cda352f985 · outbound

This paper cites A neural network approach to context- sensitive generation of conversational responses.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios A neural network approach to context- sensitive generation of conversational responses

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.692164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:32:05.039082Z digest=sha256:fedf0c2d7a29fca6e5c4902e2452d355f131f87fdd0f136e1bce3d8cd2195300

Observation 1ad2cbf5-6b74-41b7-9b27-00e041e858cf · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.053591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.053591Z digest=sha256:770293a85af6f5efd8e4f93ed830ec392505494eb4c382f9c78d643cff1dd19e

Observation 88a7abdf-79f4-43e8-81b2-56563feb7325 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.059346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.059346Z digest=sha256:dff185f49065eca7259d29b10f542d3904d132d3abb620b2d7b8eea6c9c1d6c1

Observation a5841562-877b-45c3-bd93-eb0892b3e020 · outbound

This paper cites E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.068528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.068528Z digest=sha256:36d5ae00cb6074c7ef23da0e920b12a0ab69feba5e7d05458c23110b633ab4b2

Observation e43f5829-5526-46f4-9607-2323038d955b · outbound

This paper cites AIR-bench: Benchmarking large audio-language models via generative comprehension.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios AIR-bench: Benchmarking large audio-language models via generative comprehension

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.661560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:32:05.076706Z digest=sha256:6fc390859582ee254f89b2e32ac81fd3de213de7d688c6a8d509e582cafcefdf

Observation adc1fbc1-a86d-44fb-aee8-b0fc8579a74b · outbound

This paper cites URL https://aclanthology.org/2024.acl-long.109.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios URL https://aclanthology.org/2024.acl-long.109

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.634168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:32:05.086658Z digest=sha256:ecc2e4e38a1db4ca06b39cb85f37e6bb7285221d332a4fa2a86056ff6176ca5e

Observation 01ab8409-5a99-4ad8-afe2-e538f4505f41 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios BERTScore: Evaluating Text Generation with BERT

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.104984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.104984Z digest=sha256:e51340dd6687f1b95e2bfe683cc715f38bba475be49641f4c17176e63910cbc4

Observation 877814fe-560f-402c-b46d-27869174b329 · outbound

This paper cites The cocktail fork problem: Three-stem audio separation for real-world soundtracks.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios The cocktail fork problem: Three-stem audio separation for real-world soundtracks

Reference 2002

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.769732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:32:05.003775Z digest=sha256:e443e6eee0f4f27cf913175cceb0e960a2c79fdfff098523b511679bce3e3e6a

Observation 5dc1f954-88bb-4038-835a-2cee7b608670 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.935843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.935843Z digest=sha256:9a7c0668a8c59e3fc3e59628fdbb76fbd92e5ba7b04917ff9bc972ca7da9fee3

Observation 52a13339-d729-401b-a598-487c74ecd778 · outbound

This paper cites EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.906306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.906306Z digest=sha256:73f9726011f8af5dc35a94cff4757e20e243507c81aa209aca985e5607143dd3

Observation 6c7903c7-8257-44d3-9781-757f4e9e1fec · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Audiocaps: Generating captions for audios in the wild

Reference 2009

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.808751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:32:04.958676Z digest=sha256:879ce3b2ea461e95cb140858ec3bfee8b050b207c63679ddf4ae7037e415a96e

Observation d88343bf-aec6-4c1d-8fe9-89664e868fb1 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.045777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.045777Z digest=sha256:9c10d8ea8a359384ed02a537550683b90c36cbe529548cee75abb20635111bd6

Observation f3b5d209-1f93-46e1-82d0-411060b47e74 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.095673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.095673Z digest=sha256:cf6d5c533bed585cd4e553ce63b306b2445faa49420417c5ff0c0202a16c453b

Observation fbdf322f-519a-4dba-b972-20fc0eb0faa9 · outbound

This paper cites End-to-end task-oriented dialogue: A survey of tasks, methods, and future directions.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios End-to-end task-oriented dialogue: A survey of tasks, methods, and future directions

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.713573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:32:05.028644Z digest=sha256:de23c09266723b5a498f65146e74cf783ad3998cb4e7141bd92844516df6cd44

Observation b9183a36-597e-44dd-b039-53690bfd8af4 · outbound

This paper cites Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.966190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.966190Z digest=sha256:c3dc59e167cf55d8d85fa2bd25b6a4fad431b70c43b8231ad4d757b106de0646

Observation 80b7864c-4b9e-420f-8ed2-081b2ee94f50 · outbound

This paper cites Powerset multi-class cross entropy loss for neural speaker diariza- tion.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Powerset multi-class cross entropy loss for neural speaker diariza- tion

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.744280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T22:32:05.011324Z digest=sha256:51367763058ff92570a13d74c18b357d630bcc533cb9db9fec4f3505e012894f

Observation 988a7467-632c-4296-85ca-059b6b981f90 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.916461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.916461Z digest=sha256:ec5be472fbc10ffc87168d77d927ca5fb0a1c59ff6ccc7625e3a3df1d069ef9d

Observation 11f15a96-f792-494b-9a28-65498dc40a30 · outbound

This paper cites The Llama 3 Herd of Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.944578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.944578Z digest=sha256:b1c1a13d7c59851916b818d8c317cc70a04d16a9ba7a9d8fc5816db07011885e

Pith citing papers

Observation 98d0d316-9867-42b6-906e-0f4a5197af4f · inbound

DiscussLLM: Teaching Large Language Models When to Speak cites this paper.

DiscussLLM: Teaching Large Language Models When to Speak OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:10:42.302323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T22:09:03.740109Z digest=sha256:373c7be97e9bb8d960d121df1f5d9fd6d6fb720e1d919ddaccbc2b535bbdaa04

Observation 591a55c2-6b29-4495-81f7-6d9c3820f334 · inbound

EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs cites this paper.

EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:45:45.501676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T22:52:43.396119Z digest=sha256:edb9d2bd7e8a8fb8e8b13781cb49b62cbf0b552cd1ada675cdbd43c94737d17f

Observation 34137ee2-642b-454d-b847-092e5d8a2f04 · inbound

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search cites this paper.

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:17.001064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T15:18:17.427707Z digest=sha256:83b2fd346142030787d2bd5762a4ecafb82d4d8102429872b8127858a102ab41

Observation 166861d5-7603-435a-95a3-54759ec464b2 · inbound

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search cites this paper.

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:17:28.832711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-02T23:17:03.456746Z digest=sha256:38aab587767572f5705b80f8957d58dc42226024a656f7c475583f0886e65c66

Observation 719381de-8387-42fb-8ead-4d789199210c · inbound

Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems cites this paper.

Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:16:53.731741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T04:17:18.483765Z digest=sha256:03265bc92ce3dd98efa98f0dfe5a5676a4c3595232412ca8cbd1044080e29830