Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T11:50:03.947364Z
Paper Citation Record · LEDGER
As of 1 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2606.10864.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T11:50:03.947364Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d1b14169-4b35-4c45-8544-e5f89945c72f · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Hybrid CTC/Attention architecture for end-to- end speech recognition,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b01d563-b780-4530-b8d2-e70c5c4481db · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Robust speech recognition via large-scale weak supervision,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1db1406-605b-43c8-a94d-8a2c6faa10cd · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Reproducing Whisper-style training using an open- source toolkit and publicly available data,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 669d07a2-23b5-487f-ad7d-b8b78af9ab28 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Wav2vec 2.0: A framework for self-supervised learning of speech representations,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 986f3fd8-9c41-46e4-9dbb-f71a4806f83d · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04e42d85-8d14-46d5-bd4a-9f700843e86b · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Language models are few-shot learners,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 992fab96-0548-4444-932c-5eac8b18f3ab · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition ICLEval: Evaluating in-context learning ability of large language models,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b6710c4-8ada-4492-82e0-4abca001297a · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition AudioChatLlama: Towards general-purpose speech abilities for LLMs,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a2f510a-899c-42b2-a92a-38338f8b3f75 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Prompting large language models with speech recognition abilities,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddf5143e-c500-4292-ac28-174acafe8bff · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition SALMONN: Towards generic hearing abilities for large language models,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53be9f38-2b91-475f-b8bc-f9a32fcf2d5a · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition On decoder-only architecture for speech-to-text and large language model integration,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5656288-7fba-4200-8c70-4f65b25316db · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Wav2Prompt: End-to-end speech prompt learning and task-based fine-tuning for text-based LLMs,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b8865ce-2074-48db-8f6e-c578995019b6 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Speech ReaLLM: Real-time speech recognition with multimodal language models by teaching the flow of time,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09d58f6b-d535-4145-87bc-754bb56ae523 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition SALM: Speech-augmented language model with in-context learning for speech recognition and translation,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93153887-f10e-4eb3-8b8e-95ff650fa62e · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Investigating decoder-only large language models for speech-to-text translation,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e5de177-0799-46ec-a724-2e35eb519400 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Prompting large language models with audio for general-purpose speech summarization,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df4b2ca7-7ecd-4176-ad0d-0b21386aaf62 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Contextual biasing speech recognition in speech- enhanced large language model,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36704c65-2420-4a0c-a8cb-ec950f86c00f · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition MaLa-ASR: Multimedia-assisted LLM-based ASR,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a54408b2-a99b-4b86-a9e1-2c66c8191eb1 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 284b0de5-918a-404f-9c31-6d890f22f798 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Using large language model for end-to-end Chinese ASR and NER,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83509249-a515-4525-b1cf-39a88a4c44f4 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Whispering LLaMA: A cross-modal generative error correction framework for speech recognition,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 308dcc7e-753a-4e69-9279-90908678973a · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition SLM: Bridge the thin gap between speech and text foundation models,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ef9b052-c03a-4603-bc7b-067531d782cd · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Connecting speech encoder and large language model for ASR,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8104caaf-e607-4bbb-bff9-2edd8952fe6c · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Speech recognition meets large language model: Benchmarking, models, and exploration,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2118fe69-6661-4f4f-b71d-3f4584908b88 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Improving spoken language modeling with phoneme classification: A simple fine-tuning approach,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd7db311-2d99-4413-ae0a-4b4b24da9145 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Improving large-scale deep biasing with phoneme features and text-only data in streaming transducer,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6297b633-bdeb-4f11-8f8a-7d709a229f9b · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Phoneme-aware encoding for prefix-tree-based contextual asr,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf51188d-07bb-454e-b1ee-08920e00827c · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition WavLLM: Towards robust and adaptive speech large language model,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbc33d70-9477-4bdd-915f-5130d293ceef · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition AlignFormer: Modality matching can achieve better zero-shot instruction-following speech-LLM,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab3def8d-3cf8-419b-8428-2fee0b2640fe · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition QLoRA: Efficient finetuning of quantized LLMs,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f33e143b-118a-4a18-a01b-8699a69ff287 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Speech model pre-training for end-to-end spoken language understanding,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 824d3fce-3840-4767-a90b-3331137371d9 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition TED-LIUM 3: Twice as much data and corpus repartition for experiments on speaker adaptation,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58fa854f-d45e-42d8-bd71-f1277dff0ee1 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition TED-LIUM: an automatic speech recognition dedicated corpus,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1978b4ab-abaf-4dca-902c-bbdbccba3c10 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Layer-wise analysis of a self-supervised speech representation model,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15f6a2eb-8be1-4f7e-ad30-5e48f197f1a6 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition SUPERB: Speech processing universal PERfor- mance benchmark,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4d18cae-9b84-4b3c-aad9-0b18a3a40c0d · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition The Spoken Dutch Corpus: Overview and first evaluation,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b79553a2-623e-482b-838a-fe44156eeb26 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Automatic generation of phonetic transcriptions for large speech corpora,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28860d14-9775-4695-8545-8b7daa61a088 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Assessing manually corrected broad phonetic transcriptions in the spoken Dutch corpus,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1689d178-78d0-42b0-b4b4-48f29e61a808 · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Leveraging Broadcast Media Subtitle Transcripts for Automatic Speech Recognition and Subtitling
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 0fc1c09f-bcf6-4dd7-926b-0975430163ce · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Trans-tokenization and cross-lingual vocabulary transfers: language adaptation of LLMs for low-resource NLP,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e9ebee1-5cad-42cd-8cab-f9446785a2fe · outbound
Phoneme-First Prediction for LLM-Based Speech Recognition Montreal Forced Aligner: Trainable text- speech alignment using Kaldi,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.