Pith. sign in

Paper Citation Record · LEDGER

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

As of 22 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 5 inbound Pith citation observations for arXiv:2501.01384.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.01384 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:32:05.104984Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T23:17:03.456746Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:17:28.830734Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6f033cc8-adac-4012-aae0-4cf27578d44c · outbound

This paper cites SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond Words

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.897175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.897175Z digest=sha256:0befb7105eb9e1d13f333d78d0888309fcc6976412dcafb15af60dfe10f9aa08

Observation 0d6982a2-999e-4f38-8807-0b918d657330 · outbound

This paper cites Qwen2-Audio Technical Report.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Qwen2-Audio Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.927050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.927050Z digest=sha256:a7a4c2071162c705ac395f6940ddc15166e3842fcaeeaba99a3d2e9968f4f04f

Observation ac489d20-9269-4d47-9b1f-015ee9aef61b · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios WavChat: A Survey of Spoken Dialogue Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.952630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.952630Z digest=sha256:f909009844ed64ad861a35ae5eab55488bd3b21fc2e23a5e1940a7a10d3f8a4f

Observation eea48f33-c7b0-4afa-98f3-7c95a493b137 · outbound

This paper cites Dailytalk: Spoken dialogue dataset for conversational text-to-speech.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Dailytalk: Spoken dialogue dataset for conversational text-to-speech

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.788900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:32:04.973121Z digest=sha256:48d2d3fa6d78dd19eb9d77f37125f234456e384045d56b70a462f01f13154989

Observation a1bda65c-69da-42a3-9bee-ba9a7428ecda · outbound

This paper cites Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.980466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.980466Z digest=sha256:98cb3dafc65ae537fbeabd4002ee58f1fd9f53ffedaa2bdaaed424e4640a9007

Observation cab6075e-e96e-4126-8053-974c7a93de7b · outbound

This paper cites emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.987198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.987198Z digest=sha256:065dd9083dd611f94df833617a419ce52c4801077bef9734263b4449ca860459

Observation ee9a3be0-a176-45b5-bda9-b2bdf0e8327e · outbound

This paper cites EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.997721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.997721Z digest=sha256:eb331da079605ea246a70b35aecbfbe238ac8b9807f938234ab36bd8d7496b8c

Observation afe84b7f-9da2-47d1-bc17-7dabddfe33cb · outbound

This paper cites MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.020670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.020670Z digest=sha256:892741e097102751e5c272dd2446e347f579820c6d42d6561b6a5b854eb95465

Observation 9671ed34-5a79-4e7e-bbbb-c4cda352f985 · outbound

This paper cites A neural network approach to context- sensitive generation of conversational responses.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios A neural network approach to context- sensitive generation of conversational responses

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.692164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:32:05.039082Z digest=sha256:dc949cc372ab89d0a1a349116ecfc5ebd0e62dfdc0b0648e98333a78732fcfa8

Observation 1ad2cbf5-6b74-41b7-9b27-00e041e858cf · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.053591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.053591Z digest=sha256:6f72d4efdb03a3a5384184faf9857f3bf7ae9dc51bb431aef10e454790c97da5

Observation 88a7abdf-79f4-43e8-81b2-56563feb7325 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.059346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.059346Z digest=sha256:a313aad15cd4e48bce1720fb990a6f352b18bac784dca6366782eedbc459dd37

Observation a5841562-877b-45c3-bd93-eb0892b3e020 · outbound

This paper cites E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.068528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.068528Z digest=sha256:b1f24db331fa22a1566f98031f8dd9ca60e722a6307d07bc7205f45a6436d1b7

Observation e43f5829-5526-46f4-9607-2323038d955b · outbound

This paper cites AIR-bench: Benchmarking large audio-language models via generative comprehension.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios AIR-bench: Benchmarking large audio-language models via generative comprehension

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.661560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:32:05.076706Z digest=sha256:2b8790c0df27fe56415fe57aacfc2811d775e01f71ffb26cfe0f1caef1eb11a9

Observation adc1fbc1-a86d-44fb-aee8-b0fc8579a74b · outbound

This paper cites URL https://aclanthology.org/2024.acl-long.109.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios URL https://aclanthology.org/2024.acl-long.109

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.634168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:32:05.086658Z digest=sha256:ff1514e0a5b521220264735e393952b10f94373895dd6d8498423e54b569927b

Observation 01ab8409-5a99-4ad8-afe2-e538f4505f41 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios BERTScore: Evaluating Text Generation with BERT

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.104984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.104984Z digest=sha256:cff7a0d15d38d0dd590171a324553350c262f3103e3116e53be5891c8aa3ae8e

Observation 877814fe-560f-402c-b46d-27869174b329 · outbound

This paper cites The cocktail fork problem: Three-stem audio separation for real-world soundtracks.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios The cocktail fork problem: Three-stem audio separation for real-world soundtracks

Reference 2002

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.769732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:32:05.003775Z digest=sha256:acb2f4da41886ddd163f5ed15eedbd73e1d40e7c31e6f065c923d0817edc8c07

Observation 5dc1f954-88bb-4038-835a-2cee7b608670 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.935843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.935843Z digest=sha256:1ed6fc9b610c5925f13f6842649cce644b88ac114c0475379e754bb6323575ad

Observation 52a13339-d729-401b-a598-487c74ecd778 · outbound

This paper cites EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.906306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.906306Z digest=sha256:973ba01307d18b607d3208f36d37072ca76cc1d6e36c3c23bb27a6a682c31c87

Observation 6c7903c7-8257-44d3-9781-757f4e9e1fec · outbound

This paper cites Audiocaps: Generating captions for audios in the wild.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Audiocaps: Generating captions for audios in the wild

Reference 2009

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.808751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:32:04.958676Z digest=sha256:92be01a39510efdb546d867e1f664e68cfe3520de9fc7c058238a3dfe245cbb7

Observation d88343bf-aec6-4c1d-8fe9-89664e868fb1 · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.045777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.045777Z digest=sha256:1a0fd48cf8623e9461d90acfe7bda902e292735c9474a60c220605f4460d1885

Observation f3b5d209-1f93-46e1-82d0-411060b47e74 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:05.095673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:05.095673Z digest=sha256:6d1bda8a4dbc34a4f8a1e7b79e55c570bf0dfac109bf096a81e3f7913680b741

Observation fbdf322f-519a-4dba-b972-20fc0eb0faa9 · outbound

This paper cites End-to-end task-oriented dialogue: A survey of tasks, methods, and future directions.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios End-to-end task-oriented dialogue: A survey of tasks, methods, and future directions

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.713573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:32:05.028644Z digest=sha256:4bac38d2065aeed177423eba4fd13b5da6c5ae80442ccf6fc0199ea5d8bb1354

Observation b9183a36-597e-44dd-b039-53690bfd8af4 · outbound

This paper cites Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.966190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.966190Z digest=sha256:88d62c2a92d2bef8dc55caab2b9c200779e896d679a266fae2d6fec593c150bf

Observation 80b7864c-4b9e-420f-8ed2-081b2ee94f50 · outbound

This paper cites Powerset multi-class cross entropy loss for neural speaker diariza- tion.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Powerset multi-class cross entropy loss for neural speaker diariza- tion

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:32:05.744280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:32:05.011324Z digest=sha256:69c3cd80844c3c4e35779bc622438567862b077a3cec647c2196e83079681858

Observation 988a7467-632c-4296-85ca-059b6b981f90 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.916461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.916461Z digest=sha256:3d81037ee84a16a586889fea2c55c9f95a5ab828cfcf39331a715621722ca973

Observation 11f15a96-f792-494b-9a28-65498dc40a30 · outbound

This paper cites The Llama 3 Herd of Models.

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios The Llama 3 Herd of Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T22:32:04.944578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:32:04.944578Z digest=sha256:082ca9bb2db639c69cab2d7ffd2427d1abb91803726794e6794191284b4422a4

Pith citing papers

Observation 98d0d316-9867-42b6-906e-0f4a5197af4f · inbound

DiscussLLM: Teaching Large Language Models When to Speak cites this paper.

DiscussLLM: Teaching Large Language Models When to Speak OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:10:42.302323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T22:09:03.740109Z digest=sha256:9cb7de719fece7755a5f008ef17f32c2c95efb28b6f5c3ed98b554e0529691d8

Observation 591a55c2-6b29-4495-81f7-6d9c3820f334 · inbound

EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs cites this paper.

EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:45:45.501676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T22:52:43.396119Z digest=sha256:3715994928b47b80095bf8c2294ce378ed6bcb183daf5bee811d7ff3cb88521c

Observation 34137ee2-642b-454d-b847-092e5d8a2f04 · inbound

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search cites this paper.

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:17.001064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T15:18:17.427707Z digest=sha256:d6500e8c8f03640e5401cfa0db6fd853e6f17db80051adfd22d5781fe18a97bf

Observation 166861d5-7603-435a-95a3-54759ec464b2 · inbound

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search cites this paper.

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:17:28.832711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-02T23:17:03.456746Z digest=sha256:89f5565d93a686bd87394a1e6c651ef2774d2cd8eaf60456d13c4d11d2e1de9a

Observation 719381de-8387-42fb-8ead-4d789199210c · inbound

Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems cites this paper.

Organizational Control Layer: Governance Infrastructure at the Execution Boundary of LLM Agent Systems OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:16:53.731741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T04:17:18.483765Z digest=sha256:030a44c21fe4729e07acf93a1d7410fa194ecc5cabbe01fa4f431dd1e1f7b27a