Pith. sign in

Paper Citation Record · LEDGER

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

As of 11 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 5 inbound Pith citation observations for arXiv:2501.01384.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.01384 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:32:05.104984Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T23:17:03.456746Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:17:28.830734Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6f033cc8-adac-4012-aae0-4cf27578d44c · outbound

This paper cites SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.897175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.897175Z digest=sha256:780cb831e9ca1ed0e3ca54b0638edec319c16538c280b7426366b3615b2a718a

Observation 0d6982a2-999e-4f38-8807-0b918d657330 · outbound

This paper cites Qwen2-Audio Technical Report.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Qwen2-Audio Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.927050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.927050Z digest=sha256:29b4cdbff9445dd351a4ae19cf819fbfb51b8814f56fcc960d2c1805e55404bf

Observation ac489d20-9269-4d47-9b1f-015ee9aef61b · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios WavChat: A Survey of Spoken Dialogue Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.952630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.952630Z digest=sha256:9177a958d469dd42dd44d7c2332a552c72d42400361bf3f97f919e2f078fb4f1

Observation eea48f33-c7b0-4afa-98f3-7c95a493b137 · outbound

This paper cites Dailytalk: Spoken dialogue dataset for conversational text-to-speech.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Dailytalk: Spoken dialogue dataset for conversational text-to-speech

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.788900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:32:04.973121Z digest=sha256:d1e157e90451f488c5ce0f1d3e268b3471cad4fbde291682595907612b690abf

Observation a1bda65c-69da-42a3-9bee-ba9a7428ecda · outbound

This paper cites Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.980466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.980466Z digest=sha256:ca38c238d289518a6f09ee6302e878e376ce30002e975718fdbd6174518c0301

Observation cab6075e-e96e-4126-8053-974c7a93de7b · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.987198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.987198Z digest=sha256:d189c0a7044e2d50e4da161e46ea6b568f8d57c624995437c71eb6f0a1d11bd5

Observation ee9a3be0-a176-45b5-bda9-b2bdf0e8327e · outbound

This paper cites EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.997721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.997721Z digest=sha256:6dc4d9cf2f34b0bac615a290b92857f991b2b41c9c61afd76693aba5289ae9b0

Observation afe84b7f-9da2-47d1-bc17-7dabddfe33cb · outbound

This paper cites MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.020670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.020670Z digest=sha256:fa08ae7b0d70736216dbd8abc6c6567a424f1b40eb52c135e2c54d38fa32d9a0

Observation 9671ed34-5a79-4e7e-bbbb-c4cda352f985 · outbound

This paper cites A neural network approach to context- sensitive generation of conversational responses.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios A neural network approach to context- sensitive generation of conversational responses

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.692164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:32:05.039082Z digest=sha256:e6a2a367ddcb85ac6654bfb48d8351277d00e9408dc21948377d201fb14a9194

Observation 1ad2cbf5-6b74-41b7-9b27-00e041e858cf · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.053591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.053591Z digest=sha256:770293a85af6f5efd8e4f93ed830ec392505494eb4c382f9c78d643cff1dd19e

Observation 88a7abdf-79f4-43e8-81b2-56563feb7325 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.059346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.059346Z digest=sha256:dff185f49065eca7259d29b10f542d3904d132d3abb620b2d7b8eea6c9c1d6c1

Observation a5841562-877b-45c3-bd93-eb0892b3e020 · outbound

This paper cites E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.068528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.068528Z digest=sha256:36d5ae00cb6074c7ef23da0e920b12a0ab69feba5e7d05458c23110b633ab4b2

Observation e43f5829-5526-46f4-9607-2323038d955b · outbound

This paper cites AIR-bench: Benchmarking large audio-language models via generative comprehension.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios AIR-bench: Benchmarking large audio-language models via generative comprehension

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.661560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:32:05.076706Z digest=sha256:ebb363d106c84946836b9b02cfdbeb682fa3055420007ab240437409ea1d03d3

Observation adc1fbc1-a86d-44fb-aee8-b0fc8579a74b · outbound

This paper cites URL https://aclanthology.org/2024.acl-long.109.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios URL https://aclanthology.org/2024.acl-long.109

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.634168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:32:05.086658Z digest=sha256:9b14b7aa2aef5848a07a95331374139eb4b551b69cf2b3a4b67f50f9da14c472

Observation 01ab8409-5a99-4ad8-afe2-e538f4505f41 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios BERTScore: Evaluating Text Generation with BERT

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.104984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.104984Z digest=sha256:e51340dd6687f1b95e2bfe683cc715f38bba475be49641f4c17176e63910cbc4

Observation 877814fe-560f-402c-b46d-27869174b329 · outbound

This paper cites The cocktail fork problem: Three-stem audio separation for real-world soundtracks.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios The cocktail fork problem: Three-stem audio separation for real-world soundtracks

Reference 2002

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.769732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:32:05.003775Z digest=sha256:4af2b0c0c9f3ee3293af69da59256e4f52852c1d742eb7871a9a7d07f75b0544

Observation 5dc1f954-88bb-4038-835a-2cee7b608670 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.935843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.935843Z digest=sha256:9a7c0668a8c59e3fc3e59628fdbb76fbd92e5ba7b04917ff9bc972ca7da9fee3

Observation 52a13339-d729-401b-a598-487c74ecd778 · outbound

This paper cites EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.906306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.906306Z digest=sha256:73f9726011f8af5dc35a94cff4757e20e243507c81aa209aca985e5607143dd3

Observation 6c7903c7-8257-44d3-9781-757f4e9e1fec · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Audiocaps: Generating captions for audios in the wild

Reference 2009

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.808751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:32:04.958676Z digest=sha256:c2809bc1884b3e7401180005b3456d71738d9206079a13c6a5759bf40f68c8a2

Observation d88343bf-aec6-4c1d-8fe9-89664e868fb1 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.045777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.045777Z digest=sha256:9c10d8ea8a359384ed02a537550683b90c36cbe529548cee75abb20635111bd6

Observation f3b5d209-1f93-46e1-82d0-411060b47e74 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.095673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.095673Z digest=sha256:cf6d5c533bed585cd4e553ce63b306b2445faa49420417c5ff0c0202a16c453b

Observation fbdf322f-519a-4dba-b972-20fc0eb0faa9 · outbound

This paper cites End-to-end task-oriented dialogue: A survey of tasks, methods, and future directions.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios End-to-end task-oriented dialogue: A survey of tasks, methods, and future directions

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.713573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:32:05.028644Z digest=sha256:be71ec203776ac31e5aa65d32ff74e70a29cf2e91142e4011aebb7f2f6c92165

Observation b9183a36-597e-44dd-b039-53690bfd8af4 · outbound

This paper cites Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.966190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.966190Z digest=sha256:c3dc59e167cf55d8d85fa2bd25b6a4fad431b70c43b8231ad4d757b106de0646

Observation 80b7864c-4b9e-420f-8ed2-081b2ee94f50 · outbound

This paper cites Powerset multi-class cross entropy loss for neural speaker diariza- tion.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Powerset multi-class cross entropy loss for neural speaker diariza- tion

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.744280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:32:05.011324Z digest=sha256:23d5838090831d789281200c546b0c185d5e36e0fc43211538392c2758d8e8ce

Observation 988a7467-632c-4296-85ca-059b6b981f90 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.916461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.916461Z digest=sha256:ec5be472fbc10ffc87168d77d927ca5fb0a1c59ff6ccc7625e3a3df1d069ef9d

Observation 11f15a96-f792-494b-9a28-65498dc40a30 · outbound

This paper cites The Llama 3 Herd of Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.944578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.944578Z digest=sha256:b1c1a13d7c59851916b818d8c317cc70a04d16a9ba7a9d8fc5816db07011885e

Pith citing papers

Observation 98d0d316-9867-42b6-906e-0f4a5197af4f · inbound

DiscussLLM: Teaching Large Language Models When to Speak cites this paper.

DiscussLLM: Teaching Large Language Models When to Speak OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:10:42.302323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T22:09:03.740109Z digest=sha256:92d865995fbdcf2f3157833a98ab3960812ffed407d39ec2dbb84dc00a2668b5

Observation 591a55c2-6b29-4495-81f7-6d9c3820f334 · inbound

EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs cites this paper.

EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:45:45.501676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T22:52:43.396119Z digest=sha256:8723a04e2bb51752f310a37b419502d8274c14fd77e41b631d88d13c2ff721d1

Observation 34137ee2-642b-454d-b847-092e5d8a2f04 · inbound

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search cites this paper.

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:17.001064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T15:18:17.427707Z digest=sha256:04d42b5f851339ef2183a8135839d1367e39c1d0782cc7b77b48fd0786981019

Observation 166861d5-7603-435a-95a3-54759ec464b2 · inbound

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search cites this paper.

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:17:28.832711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-07-02T23:17:03.456746Z digest=sha256:6b809e66f502b59430e07fc5ada60bc106cacdfba9809dc26a050bb5cdac26e7

Observation 719381de-8387-42fb-8ead-4d789199210c · inbound

Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems cites this paper.

Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:16:53.731741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T04:17:18.483765Z digest=sha256:cff1fda282f2715c57d1f453dfe51fe747b209f7307a492f87657abc70351306