Pith. sign in

Paper Citation Record · LEDGER

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment

As of 7 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2506.06343.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06343 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:58:44.194819Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T05:50:33.363981Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c806618a-d343-424a-993a-e6ff54626dfd · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Moshi: a speech-text foundation model for real-time dialogue

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.491594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:40.491594Z digest=sha256:2434326c37cf8ffb22fd421a9c5b1bfaf7efb53bf6dc37ec60906720bb4fd44e

Observation cb7aac0f-dd7d-4a7b-b1c4-2e3a9c3bc669 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.561463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:40.561463Z digest=sha256:7987848d5fdd917601f83a2cf59b49a0f6d325538f141dcd2c1b9cf48a81762f

Observation 4a70dcad-3b11-496f-98e0-1fda8883fab2 · outbound

This paper cites Qwen2-Audio Technical Report.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Qwen2-Audio Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.684392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:40.684392Z digest=sha256:9212bb2778f68b00b3e0070b06822e6ee69e5abcfbbaf678dea60350160a0fd7

Observation 631d0733-5633-44b9-8519-626565e9a63e · outbound

This paper cites Baichuan-omni-1.5 technical report,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Baichuan-omni-1.5 technical report,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.810890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:40.810890Z digest=sha256:fb142e163c101268455b3473376c686b25258df592dcbeb44edf859ccbc665dc

Observation 076e0b54-13c8-48fa-aa92-1a3b5b45342a · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.924179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:40.924179Z digest=sha256:e9b0d99d6a17aeeba50b8e23b65f2f05450ba72c7f1b03c23f69de4efeca3e06

Observation 0e806ba7-755f-4f05-b107-ee54be2e91b8 · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.015033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.015033Z digest=sha256:9ed88382eaddf2f3ba1db4f34108713cf3456b6bbea8749625412807416bc0e9

Observation 123184ec-ddfa-45d5-854a-097c2d554f96 · outbound

This paper cites BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.157874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.157874Z digest=sha256:3d4972333bbabedfaacfc41ac05a1440d38472107c9c7829bf060c4f658e7be3

Observation bed4a7ab-c49e-4df2-8453-db75df57dc7e · outbound

This paper cites InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.275667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.275667Z digest=sha256:b1273695ffd1aa252f77fd7d8c495865ebc50dbbfda7015d64f7ef2a1d935f8e

Observation 3fec8c50-3c76-4ff4-837a-7c90a0fa5985 · outbound

This paper cites Distilling an End-to-End Voice Assistant Without Instruction Training Data.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Distilling an End-to-End Voice Assistant Without Instruction Training Data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.386187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.386187Z digest=sha256:43bf67f33dd4e30111fdc768165e8035a621e622285e66f29352fff0e86b32ad

Observation 6df42d5f-d314-4609-a221-5560619ae578 · outbound

This paper cites Speechless: Speech Instruction Training Without Speech for Low Resource Languages.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Speechless: Speech Instruction Training Without Speech for Low Resource Languages

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:58:44.627131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:58:41.531921Z digest=sha256:eaeaac75fb9b1c230d2e144bbb849a0a26566691b32bae1bfe5bd4660e52ccfb

Observation 4e0f11e2-22a8-4db6-94d5-26769db66329 · outbound

This paper cites SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.673699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.673699Z digest=sha256:198814dd67edceff0775639f71c13dca28994443051c67e3bc8043d6f90b41bc

Observation c6c26aea-4ace-4858-abee-3cebbce44cf4 · outbound

This paper cites SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment SLAM: A Unified Encoder for Speech and Language Modeling via Speech-Text Joint Pre-Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.805991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.805991Z digest=sha256:6ba33e5d8d2f2ef52fb13b7ce023d5c090fe0f3990fafd3ab24be217e3cdb1c4

Observation 073b2990-7cce-476c-be3a-76ec126d729b · outbound

This paper cites mSLAM: Massively multilingual joint pre-training for speech and text.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment mSLAM: Massively multilingual joint pre-training for speech and text

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.929162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.929162Z digest=sha256:bd6136e0fa266db427daa741800704c2dd0103e191ae47bf092279e0cf1ab5c1

Observation 7cc83ce2-6990-4d94-9e94-e5b4ad163fc3 · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:42.065571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:42.065571Z digest=sha256:52fb25339196c9aadb6f83652a1ae8540233a931509eb825dd8ada703c2e63eb

Observation d1f7fc13-e024-4ffa-882b-2ba265f37872 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:42.240760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:42.240760Z digest=sha256:b3c40fd60034137c693c64290fe04e4cdc013b497f7b2683f79a4fb6877d0394

Observation 4e7d927f-2eff-4969-982b-308d2f6ec336 · outbound

This paper cites W2v-bert: Combining contrastive learning and masked language mod- eling for self-supervised speech pre-training,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment W2v-bert: Combining contrastive learning and masked language mod- eling for self-supervised speech pre-training,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:42.346892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:42.346892Z digest=sha256:a83acbaa4e8e8c1cb7a9d9a69f738b999dd63e2a7af5ff408f4f64f58599e7f7

Observation a3c87180-6074-434f-bc2e-2c9d23dd2acd · outbound

This paper cites VoiceBench: Benchmarking LLM-Based Voice Assistants.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment VoiceBench: Benchmarking LLM-Based Voice Assistants

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:42.482144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:42.482144Z digest=sha256:7b3f7cf9e3fdcdf3175e03c30be7aee82339ebf86ecf26a5530776dcbccb3227

Observation 6ad31277-d93e-4068-a1cd-f9b30221ff29 · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:42.585526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:42.585526Z digest=sha256:b68da028f5381188da8456df1699cd96b14b60f4040759ebb8be3155d58a8bdb

Observation 87aa7824-7b99-471b-86b6-de5acded7abb · outbound

This paper cites Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:42.752950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:42.752950Z digest=sha256:d50b20006c606e552a382400022c62b67c58975d71f4a8ad232cf8bed3005f58

Observation f8818343-d73f-4f41-966a-38e8a1e32526 · outbound

This paper cites Spirit-lm: Interleaved spoken and written language model,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Spirit-lm: Interleaved spoken and written language model,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:46.543567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:58:42.886123Z digest=sha256:0aae5bf5310255216fb6498a05371c58f5b7aa58a5bed8a226edfa924e513d35

Observation cb7bff6c-c7a8-4840-8086-89d7b676e829 · outbound

This paper cites Dissecting learning and forgetting in language model finetuning,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Dissecting learning and forgetting in language model finetuning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:46.248933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:58:43.012070Z digest=sha256:464c2f92a309dc0950fc8ce845c62f9dd8a164a110e8967a54f30151de20d733

Observation 12c6ad8f-6c7b-484d-8d59-4e2841806b24 · outbound

This paper cites The Llama 3 Herd of Models.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment The Llama 3 Herd of Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:43.136188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:43.136188Z digest=sha256:11ebf85757402353d5658ec22f5e0e592175978fa3e3898dcd93b698d3bac2dc

Observation e3aa62cf-f587-4f1a-a40b-15fce72d20b0 · outbound

This paper cites Openwebtext corpus,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Openwebtext corpus,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.929557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:58:43.312559Z digest=sha256:5b2e58be5315d7ab9b18ab377ef5d56c6a7b631ce0646efed0ff0a7a072a4934

Observation 09a79cd4-a638-4dfc-a433-35c79d77fbd4 · outbound

This paper cites Enhancing chat language models by scaling high-quality instructional conversations,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Enhancing chat language models by scaling high-quality instructional conversations,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.639363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:58:43.450637Z digest=sha256:a8e330deb73b3beac5a216986a7301af12ddcd018352c302d54413b71d88ddda

Observation 46b02533-0f52-45e5-9613-d44a4c88cc04 · outbound

This paper cites Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Openhermes 2.5: An open dataset of synthetic data for generalist llm assistants,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.383713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:58:43.579967Z digest=sha256:f1347eafe3e1c195a57ec7c27a9d7773dd6332e2f459965d3f9447a4abcb8f6b

Observation b10ec70f-cf80-4788-8a44-f29e9e14ff3e · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Can a suit of armor conduct electricity? a new dataset for open book question answering,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:43.752110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:43.752110Z digest=sha256:1863d4af856b559a85babb261aefd718e815e67650e1cc15f7da96567913b1e2

Observation dfadb9da-0797-43a0-b1b7-f9913edbc643 · outbound

This paper cites CommonsenseQA: A question answering challenge targeting commonsense knowledge,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment CommonsenseQA: A question answering challenge targeting commonsense knowledge,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:45.100191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:58:43.896352Z digest=sha256:066164bae7b1e152e540cf7dfaf0a50a7ab65abe9298f64eee93f98da20956bb

Observation bb23e8f6-66ff-44a6-8976-98ff6d1c6755 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment Librispeech: an asr corpus based on public domain audio books,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:44.064280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:44.064280Z digest=sha256:d7ce95b848359d8592e2cc1df34a7f08195bd925db27ffdd46dea739a0d9e248

Observation 2a8231c3-cc6a-4139-afe9-79d01fc0f81a · outbound

This paper cites CoVoST 2 and Massively Multilingual Speech-to-Text Translation.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment CoVoST 2 and Massively Multilingual Speech-to-Text Translation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:44.194819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:44.194819Z digest=sha256:bead88ff670239714e4c9053c52967d2493ecded2c5663447023088807f39668

Pith citing papers

Observation 1302492f-f91d-49b9-82bd-e80a90afee7b · inbound

MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond cites this paper.

MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T05:50:33.363981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:50:33.363981Z digest=sha256:bb4e4ab100f14e7adab05fdb1a269f404d4effc3431482a18dffa2640557e04e