Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:46:16.782429Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 4 inbound Pith citation observations for arXiv:2412.01078.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:46:16.782429Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:52:16.074875Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T08:24:03.552101Z
13 of 13 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1b052b07-b8ff-4569-8727-922ce7be9484 · outbound
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7a01eef5-f9ec-4282-b3cf-a5c3ccfe2db8 · outbound
Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Qwen2-Audio Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91cf2943-42c7-470c-a3dc-3a76cd0cb803 · outbound
Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 27835fd9-0449-4821-bda1-84734b888dcf · outbound
Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83205df4-17bf-443e-9a74-7d2871d0382f · outbound
Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data MLS: A Large-Scale Multilingual Dataset for Speech Research
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aaaaffc-32d2-4273-977e-4108622727d1 · outbound
Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5efdad44-6388-4046-acd6-3c62a542dd26 · outbound
Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36698443-574c-4354-b9f7-88641718841d · outbound
Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1fdf82f8-b84e-45c2-a7b1-230f50630767 · outbound
Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d0fc32ca-08c7-43ad-97a6-c069c226c1a3 · outbound
Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 56bdf33f-6dbe-4b09-91c6-31a3af7406b5 · outbound
Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data is_suitable_for_speech
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 88578298-d84e-4d76-9086-eb5c3100786c · outbound
Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Pro- cessing (ICASSP), pages 6968–6972
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4e5744f6-03e9-4d65-8dd9-c5a8f5045879 · outbound
Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data VoiceBench: Benchmarking LLM-Based Voice Assistants
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46fcac79-c811-47f7-8aa0-85c9f5e7ab93 · inbound
DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69e3da9f-8225-4c8d-bd59-84db56e65a27 · inbound
SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f8ff0edb-4d33-4701-92c7-17ec5418c873 · inbound
Voice "Cloning" is Style Transfer Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a54bac18-3bdd-4ef9-9df4-cdabaf8bb88b · inbound
Voice "Cloning" is Style Transfer Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.