Pith. sign in

Paper Citation Record · LEDGER

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data

As of 20 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 4 inbound Pith citation observations for arXiv:2412.01078.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01078 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:46:16.782429Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:52:16.074875Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T08:24:03.552101Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1b052b07-b8ff-4569-8727-922ce7be9484 · outbound

This paper cites this" or.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data this" or

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:16.985455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T04:46:16.760093Z digest=sha256:d302d90a570ebf553c05c583d5dc93e05fa34571624d6844a715a192561358f9

Observation 7a01eef5-f9ec-4282-b3cf-a5c3ccfe2db8 · outbound

This paper cites Qwen2-Audio Technical Report.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Qwen2-Audio Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:16.729612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:16.729612Z digest=sha256:8a966fe550621e1722f061b984dca1cb882f2e715ba630321bb502001151b2e4

Observation 91cf2943-42c7-470c-a3dc-3a76cd0cb803 · outbound

This paper cites an unresolved cited work.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:46:16.953950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T04:46:16.769330Z digest=sha256:25b1da2605022e42a0ea7b4db142a8f6e4e4c99f40096b7aefead00651d92b91

Observation 27835fd9-0449-4821-bda1-84734b888dcf · outbound

This paper cites Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:16.739240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:16.739240Z digest=sha256:38fe313973272b04c8a0cfc9915720702d69e3fa92212f6900d5ce4459718394

Observation 83205df4-17bf-443e-9a74-7d2871d0382f · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data MLS: A Large-Scale Multilingual Dataset for Speech Research

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:16.744798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:16.744798Z digest=sha256:c2d09ef782e82b25ab1036d021fce8eea5b8b7021e1a3f73a7712d3c37b5e57e

Observation 5aaaaffc-32d2-4273-977e-4108622727d1 · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:16.749983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:16.749983Z digest=sha256:bcf82d5a17a789c58266bf5041167483326ffa17ad9b61079ee889e5a9eaaa47

Observation 5efdad44-6388-4046-acd6-3c62a542dd26 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:16.755242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:16.755242Z digest=sha256:ac26e5c90e005f9bc3d4d2c99ad982ebc98ca393f0a987490585a36b0e07060e

Observation 36698443-574c-4354-b9f7-88641718841d · outbound

This paper cites an unresolved cited work.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:46:16.968375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T04:46:16.764723Z digest=sha256:b358252a24311bf128b67ec49789a6c93bbdf8dbb1640b1977b976e160714532

Observation 1fdf82f8-b84e-45c2-a7b1-230f50630767 · outbound

This paper cites an unresolved cited work.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:46:16.939317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T04:46:16.774097Z digest=sha256:ed5d07ca0ebb43c8c2e0079b06ce5e5ba0f51150981cfb28fbe9bcb545522083

Observation d0fc32ca-08c7-43ad-97a6-c069c226c1a3 · outbound

This paper cites an unresolved cited work.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:46:16.924599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T04:46:16.778156Z digest=sha256:2b9da98be38cb29e8ef7f96307db8046b210cf6c8d6383f6acfe68243b1ddcd5

Observation 56bdf33f-6dbe-4b09-91c6-31a3af7406b5 · outbound

This paper cites is_suitable_for_speech.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data is_suitable_for_speech

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:16.909892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T04:46:16.782429Z digest=sha256:1a753c1cb97b71373258c7ac30eb49e9eee406052be0aa084ba91ef62d5da32b

Observation 88578298-d84e-4d76-9086-eb5c3100786c · outbound

This paper cites In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Pro- cessing (ICASSP), pages 6968–6972.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Pro- cessing (ICASSP), pages 6968–6972

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:46:17.001536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T04:46:16.734607Z digest=sha256:aea0e1c1d34c3f2c1aadc18ec47fdf1f4f6970241bbfbc6ac88ba125d69d8c4f

Observation 4e5744f6-03e9-4d65-8dd9-c5a8f5045879 · outbound

This paper cites VoiceBench: Benchmarking LLM-Based Voice Assistants.

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data VoiceBench: Benchmarking LLM-Based Voice Assistants

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T04:46:16.724250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:46:16.724250Z digest=sha256:0c91d3c0afa542a7367f1400b642abfcb7e3ce6f630d32323119bc7d3a83b276

Pith citing papers

Observation 46fcac79-c811-47f7-8aa0-85c9f5e7ab93 · inbound

DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue cites this paper.

DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:52:16.074875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:52:16.074875Z digest=sha256:633d39a6566920b91e75f054dd920062d392df918b7fe68dcfd673a94578c4f7

Observation 69e3da9f-8225-4c8d-bd59-84db56e65a27 · inbound

SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation cites this paper.

SpeechMedAssist: Efficiently and Effectively Adapting Speech Language Models for Medical Consultation Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T17:01:07.569586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T17:00:59.421441Z digest=sha256:de616129cd06b746c9c6fdc3126dda4f860a9f1bbc1a2393df247a9b77e1e688

Observation f8ff0edb-4d33-4701-92c7-17ec5418c873 · inbound

Voice "Cloning" is Style Transfer cites this paper.

Voice "Cloning" is Style Transfer Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:02:47.145012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-19T21:01:24.307137Z digest=sha256:15a8c34ceeb1d8d2fbbe0832d938a2ae435368d4c08d6e09d13b38015d6ece3c

Observation a54bac18-3bdd-4ef9-9df4-cdabaf8bb88b · inbound

Voice "Cloning" is Style Transfer cites this paper.

Voice "Cloning" is Style Transfer Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:24:03.553542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-21T08:20:29.918036Z digest=sha256:dcaef2154d9f02e979d1669d35e11009bfa93bebe6adbb253199f2d893cf3f65