Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:44:25.985859Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2412.01145.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T04:44:25.985859Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-22T20:44:57.476464Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-22T20:45:08.147897Z
48 of 48 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 985a90aa-debf-49ca-be33-294fffdccfa5 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Language models are few-shot learners,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 716d7f8d-3a7c-4636-b8df-37eb5a2e534d · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cb7d0de-fd21-46ad-abff-27c6cbd21103 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM The Llama 3 Herd of Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8df4194-3025-4b92-abd6-a0202eeb08ec · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Self-instruct: Aligning language models with self- generated instructions,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3cdce924-0448-4e9e-bce3-0b7308d2b00e · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Training language models to follow instructions with human feedback,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 329c671c-df29-4ec5-92ad-516a771e5422 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Direct preference optimization: Your language model is secretly a reward model,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a2cc399-9573-4da9-aefa-2645b242341f · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Moshi: a speech-text foundation model for real-time dialogue
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06b2145c-2d85-4191-bde8-1c5a2f153906 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36998284-953c-477a-95da-df47762a1d60 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c733e4d-e0fc-41df-9404-1b8e19304a9a · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19c436fa-ff2f-4d68-8e51-6a9e950764e1 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fedfc8f9-0c10-4bc1-80ab-eae0ffdfc593 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e60d1685-300a-4a71-b8d3-9d17e046fbbc · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4f4b24e-79ee-4853-945e-ef4672dfdf7d · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SALMONN: towards generic hearing abilities for large language models,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5997242c-74a7-4943-81cf-1c028e5d6d06 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Qwen2-Audio Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 181cedd0-34f3-49bf-a163-cc409242e82e · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM On decoder-only architecture for speech- to-text and large language model integration,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 782fcbb5-d293-44d4-b352-6e7db367d78f · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM COSMIC: data efficient instruction-tuning for speech in-context learn- ing,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e622f2e9-4fe5-42dc-bd14-35fab6d28c79 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Prompting large language models with speech recognition abilities,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7df69f19-523c-44be-8936-bd2a9a07ecf9 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SpeechVerse: A Large-scale Generalizable Audio Language Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97d59d3c-1b84-4e41-85d6-add42654e0e4 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM High fidelity neural audio compression,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 623da227-e956-4251-886e-484a7e084a54 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM High- fidelity audio compression with improved rvqgan,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 818a281e-6dc3-40a4-a1bd-5befccdef7f2 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 873c4ec5-451a-4960-a1e5-be8b46640b66 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Wavllm: Towards robust and adaptive speech large language model,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation da070468-5174-4080-9257-9d6e12c61ff6 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Audiochatllama: Towards general-purpose speech abilities for llms,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b6e663c6-107c-432c-9304-37c183a30bc2 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44f31491-4a10-408e-9d3f-60af5e18d21c · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36749f60-17f5-410b-810f-9da282938ede · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c5b8b81-e1a0-439a-b105-2433776bea43 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Wav2Prompt: End-to-end speech prompt learning and task-based fine-tuning for text-based LLMs,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 92414057-f770-43da-ad24-19aa9404c3a0 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36333554-2450-44f8-8090-6b818c617999 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65626896-ee5e-4025-a331-63ae60dc2ac7 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Connec- tionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 71511e59-350c-4d5c-a4fd-e8c8a0a96a34 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM CASS-NAT: CTC alignment- based single step non-autoregressive transformer for speech recognition,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e5e6650f-2ee5-4b1c-ab49-d42a5d3749a0 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Unienc-cassnat: An encoder-only non-autoregressive asr for speech ssl models,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9c2802fb-909d-4e2d-b02f-3ac59ab22d71 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Ctc-based compression for direct speech translation,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 982d628b-f3e8-4a9b-af47-f524d73fcd9d · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a2d75a61-cb1c-437a-bc2b-b126ea13604c · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SpeechT5: Unified-modal encoder-decoder pre-training for spoken language processing,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 85514468-a283-4725-a254-7017d248630d · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SpeechLM: Enhanced speech pre-training with unpaired textual data,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 86f05e72-8f1e-4a42-a720-c8b46a45e795 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Seamless: Multilingual Expressive and Streaming Speech Translation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e74b7024-a6db-44f6-98ec-f26174eec041 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM M-adapter: Modality adaptation for end-to-end speech-to-text translation,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c0d740d2-39c1-4905-b8be-c793fb4494d9 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM MAESTRO: Matched speech text representa- tions through modality matching,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c00317ff-d02d-48b0-9d02-c6ad41888208 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Cjst: Ctc compressor based joint speech and text training for decoder-only asr,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3b3ba2a4-aa92-49ee-829b-85887d30e4fd · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Lora: Low-rank adaptation of large language models,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7a096d0b-5c63-45e5-9e99-4ca0b139bb8c · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8c2d6d40-f604-424b-abfb-076846be80d0 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Conformer: Convolution-augmented transformer for speech recognition,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d6c1a48e-e5b1-4d07-9ecf-d18ee2837067 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM A CTC alignment-based non- autoregressive transformer for end-to-end automatic speech recognition,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 490ff960-f90d-44d6-a8f0-fde17f19def0 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Unsu- pervised cross-lingual representation learning at scale,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4e3c3e89-b628-42c0-b2c0-3cfcd5f1b980 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Zero: Memory optimizations toward training trillion parameter models,
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a11f7ff-e3ee-4275-81e7-b331e95a2782 · outbound
AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM 4552–4572
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 365b54bc-1661-477d-8383-51bc92147390 · inbound
On The Landscape of Spoken Language Models: A Comprehensive Survey AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.