Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T11:32:58.335783Z
Paper Citation Record · LEDGER
As of 1 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2606.10853.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-27T11:32:58.335783Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-27T11:32:58.335783Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T07:57:44.678976Z
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 60fed204-647c-4b3c-9f12-58151d5d53f8 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6f9d7b5-ca60-4e2a-bb2f-3f8d3fe1937f · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Speech Encoder Fusion for LLM-based Automatic Speech Recognition
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 010a30db-6313-4ea5-9202-5a6d35e625d6 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89867cb9-6195-48ab-9ba3-1f4894f17068 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35356ebf-7ce3-40e3-b0e7-3e86e9e3077e · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb5908c9-155e-46a0-b3a7-775e4b9a98b1 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition 0: Hello. 1: How are you? 0: I’m fine
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ada9d18-f7dc-4a05-b059-d6ee9e61bb5c · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Although speech-LLMs have difficulties at- taining performance of dedicated ASR systems, augmenting LLMs with speech capabilities has much wider applications be- sides ASR
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a87ba39-287d-41aa-8460-838167b111fa · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition We found that careful fusion outperforms standard feature concatenation in all cases
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b477bcbc-5e85-4293-bb4c-730f1b97f521 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6049178e-9f83-4431-92f0-7f21f52becfd · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition No part of the manuscript’s content or ideas was produced by generative AI
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 487ef033-f0e3-49fc-b85c-7f06de76871a · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition BLIP-2: Bootstrap- ping language-image pre-training with frozen image encoders and large language models,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bb29882-95e5-4591-8d44-620d6e41f88a · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Visual instruction tuning,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4925622-4525-4f07-9bf9-2d1dc972c69e · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 9603f22e-0b20-4bfc-9826-3f197c8a8bef · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition SALMONN: Towards generic hearing abilities for large language models,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e103e6f-5f4c-4c24-b3fe-00e35ea8fb3b · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition AudioChatLlama: Towards general-purpose speech abilities for LLMs,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdd62753-c608-423c-990c-2ba71b132fac · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Speech ReaLLM – real-time speech recognition with multimodal lan- guage models by teaching the flow of time,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 318b608d-a294-4525-984e-d6a297799777 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Wav2Prompt: End-to-end speech prompt learning and task-based fine-tuning for text-based LLMs,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff5e17ce-a7c4-4af2-92ab-3524031ea049 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition BESTOW: Efficient and streamable speech language model with the best of two worlds in GPT and T5,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46b05289-1338-4f5b-b589-5d650f88aafe · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Speech recognition meets large language model: benchmarking, models, and exploration,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bfff35c-5238-401a-8381-ced8bbedb4c5 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Connecting speech encoder and large language model for ASR,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f4fbd24-82ab-41d3-a5b8-df67363b844d · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition How to connect speech foundation models and large language models? What matters and what does not,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fab720ae-6a5c-40a0-adbc-f91b6e861fc4 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Prompting large language models with speech recog- nition abilities,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50653965-3851-4867-a341-582a0d4625e0 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Robust speech recognition via large-scale weak su- pervision,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b0887a6-ae9c-43fa-910f-679b66637a3c · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition OWSM v4: Improving open whisper-style speech models via data scaling and cleaning,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c8523ea-90b5-4524-af70-4043a0286dc0 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0698eaf9-bb2a-4352-bfd4-1d80d95273f2 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition WavLM: Large-scale self-supervised pre-training for full stack speech processing,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82620119-0285-45c0-819e-13e01dbf8397 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Better pseudo- labeling with multi-ASR fusion and error correction by speech- LLM,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 962645b5-c3a8-4b21-8318-6fe2577589a6 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Attention is all you need,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6009c95-e63b-4404-a6a4-4bfbfe68c066 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Leveraging broadcast media sub- title transcripts for automatic speech recognition and subtitling,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51d583c8-c661-4845-81b7-2939c33c77aa · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition WavLLM: Towards robust and adaptive speech large language model,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 491b754c-fb6f-4348-9336-9153c19c506a · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Diarization-Aware Multi-Speaker Automatic Speech Recognition via Large Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 802cb44b-d06e-405c-84fa-d60e55921e2b · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition MoWE-Audio: Multitask audioLLMs with mixture of weak encoders,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43987fa7-989d-40d4-a183-53a5ca314ad3 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Enhancing speech large language models with prompt-aware mixture of audio encoders,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a2344da-645b-48b6-9468-16787507174e · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Hybrid CTC/Attention architecture for end-to-end speech recog- nition,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9693f9bd-1ffd-4fcf-8157-511f9c630a2c · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Learning to jointly transcribe and subtitle for end-to-end spontaneous speech recognition,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e4501f9-0825-450a-a54e-3f2ef6acbc6f · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Lib- rispeech: An ASR corpus based on public domain audio books,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab6a9c0b-3afe-40e0-9150-5e98d0008ff2 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition ECAPA2: a hybrid neural net- work architecture and training strategy for robust speaker embed- dings,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 464ba1d2-b0fb-4563-bf84-daaabbcea227 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Trans-tokenization and cross-lingual vocabulary transfers: language adaptation of LLMs for low-resource NLP,
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feffa023-88b5-4fd3-88c2-4af5cbdad417 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition AlignFormer: Modality matching can achieve better zero-shot instruction- following speech-LLM,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8f7c0ef-d25f-47f8-8e0a-65a1d4622628 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition The spoken Dutch corpus. overview and first evalu- ation,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66106ef6-4cfd-485c-be69-68355538d9b7 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Seri- alized output training for end-to-end overlapped speech recogni- tion,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39428aff-2077-43df-917f-e697a134724e · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition The rich transcription 2007 meeting recognition evaluation,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dfbb77c-492b-4d08-887e-995a839d3e90 · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual Speech Recognition Evaluation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation 3c489005-8d64-402a-a936-602886a09bba · outbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Evaluation of LLMs in speech is often flawed: Test set contamination in large language models for speech recognition,
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6f9d7b5-ca60-4e2a-bb2f-3f8d3fe1937f · inbound
Speech Encoder Fusion for LLM-based Automatic Speech Recognition Speech Encoder Fusion for LLM-based Automatic Speech Recognition
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.