Pith. sign in

Paper Citation Record · LEDGER

Distilling an End-to-End Voice Assistant Without Instruction Training Data

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2410.02678.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.02678 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:27:27.223349Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T20:37:34.403610Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 08370aae-ac17-40f6-9ace-641da07db20c · inbound

VoiceBench: Benchmarking LLM-Based Voice Assistants cites this paper.

VoiceBench: Benchmarking LLM-Based Voice Assistants Distilling an End-to-End Voice Assistant Without Instruction Training Data

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:50:13.980645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-17T00:50:13.841689Z digest=sha256:7b5c831b54740b918531b5de00105f09ed54e107695df24fc89cb5a3aa572ff4

Observation 91cc5029-1500-4b98-b8c1-4dd268680a26 · inbound

SpeechPrune: Context-aware Token Pruning for Speech Information Retrieval cites this paper.

SpeechPrune: Context-aware Token Pruning for Speech Information Retrieval Distilling an End-to-End Voice Assistant Without Instruction Training Data

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T14:27:27.223349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:27:27.223349Z digest=sha256:66d79771251c77b592caa63717cce93027b9bbd3c11e6f2d5d820f85555513ac

Observation fdab8341-e8c9-4432-9e39-391f534c8a06 · inbound

Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models cites this paper.

Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models Distilling an End-to-End Voice Assistant Without Instruction Training Data

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T12:56:13.139774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:56:13.139774Z digest=sha256:547b33ccf5f4c009df6d145470a0c33bef252e1d79dbf4577b867636bac32b0b

Observation 0d9e04fd-c1fd-4b44-95e7-e2073d6d9188 · inbound

Contrastive Learning for Task-Independent SpeechLLM-Pretraining cites this paper.

Contrastive Learning for Task-Independent SpeechLLM-Pretraining Distilling an End-to-End Voice Assistant Without Instruction Training Data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:13:48.475360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:13:48.475360Z digest=sha256:695e5127b194f0677313ada38cd151785f6ad98368face31af8307f6aa57943e

Observation 1d580a2c-1d04-4963-96aa-98457ad445ca · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey Distilling an End-to-End Voice Assistant Without Instruction Training Data

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T20:45:08.182294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:7917fe2f7065badf5afeae12cc8b9f893bd948fb01f7177010bb31876102aab1

Observation e9bebd05-8fbe-4d53-b74d-843f9e9c173f · inbound

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models cites this paper.

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models Distilling an End-to-End Voice Assistant Without Instruction Training Data

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:01.636740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:23:01.636740Z digest=sha256:d950a1dd009217089ad18911165c369d908eedd68835d4a9b6a052e077d5c160

Observation caa0ee6b-f3a4-45e4-98b1-de8f6327de53 · inbound

Speechless: Speech Instruction Training Without Speech for Low Resource Languages cites this paper.

Speechless: Speech Instruction Training Without Speech for Low Resource Languages Distilling an End-to-End Voice Assistant Without Instruction Training Data

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:17.014767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:17.014767Z digest=sha256:c637f419afea53bad183a4d293e31f5bafb9999008503bfe77ca762556ca0d65

Observation 9b55b015-bacf-4433-979d-63512062722d · inbound

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models cites this paper.

Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models Distilling an End-to-End Voice Assistant Without Instruction Training Data

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:33.963171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:33.963171Z digest=sha256:64bb6aa860c3115f24c8053c645284d8bbcbbdc2b33fb34c9c23f620864854e0

Observation e9c68199-96f5-4c3a-8ce6-0ca1a4b2e412 · inbound

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems cites this paper.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems Distilling an End-to-End Voice Assistant Without Instruction Training Data

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:31.596425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:31.596425Z digest=sha256:2284d81b4f334592952709820eb468d720b18403951f6cd755356eb8fdedbe15

Observation 3fec8c50-3c76-4ff4-837a-7c90a0fa5985 · inbound

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment cites this paper.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Distilling an End-to-End Voice Assistant Without Instruction Training Data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.386187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.386187Z digest=sha256:397efb4ca0554494264ef24d8746d10fa0541383e74764d9952ad98b0b376e51

Observation 7fb330a8-2c7b-4421-a33b-10711326c2a2 · inbound

AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation cites this paper.

AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation Distilling an End-to-End Voice Assistant Without Instruction Training Data

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T16:47:55.927357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:47:55.927357Z digest=sha256:a4a87cc1f56faf9058a30a5e7267e4dc2f78c87a91404d27f0f0c43483a40e3b

Observation 5ab66f10-a6c1-4589-a7d7-238edc5ca84d · inbound

Benchmarking Gaslighting Attacks Against Speech Large Language Models cites this paper.

Benchmarking Gaslighting Attacks Against Speech Large Language Models Distilling an End-to-End Voice Assistant Without Instruction Training Data

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:10:31.487105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T08:06:59.951859Z digest=sha256:1cb08f44a55ea813e7d20046f72a6c50fb44bd99193dc9beafbfbdc1800e3e0e

Observation 14c54e88-8246-4c65-b344-c9b593196b10 · inbound

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs cites this paper.

Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs Distilling an End-to-End Voice Assistant Without Instruction Training Data

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-10T20:37:34.405190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-10T20:30:11.127007Z digest=sha256:1c04b1555fef27ebe6c5ed5868e9e4bf1e0ceab96f15028ebab8d84fdff7afc8