Pith. sign in

Paper Citation Record · LEDGER

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM

As of 13 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2412.01145.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01145 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:44:25.985859Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T20:44:57.476464Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T20:45:08.147897Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 985a90aa-debf-49ca-be33-294fffdccfa5 · outbound

This paper cites Language models are few-shot learners,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Language models are few-shot learners,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.697786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.784371Z digest=sha256:c57da95aa74654dc1ce4fc8fc804d0616a44158eb54784a9447cc4141cfc3ed0

Observation 716d7f8d-3a7c-4636-b8df-37eb5a2e534d · outbound

This paper cites GPT-4 Technical Report.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.790256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.790256Z digest=sha256:2821f93a1368702f90d9189570303d49361fe6b4763f2e3cfa9c770fb6fdc8bb

Observation 0cb7d0de-fd21-46ad-abff-27c6cbd21103 · outbound

This paper cites The Llama 3 Herd of Models.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.795189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.795189Z digest=sha256:6fb177618e3cf4b67246a45e4744727e37f84a88b117b531f83ee806d992d835

Observation c8df4194-3025-4b92-abd6-a0202eeb08ec · outbound

This paper cites Self-instruct: Aligning language models with self- generated instructions,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Self-instruct: Aligning language models with self- generated instructions,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.680275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.801322Z digest=sha256:8b433624d699106e5e5c45340fac78bddbe2ca68429e4ea4f71ba553662ce6a7

Observation 3cdce924-0448-4e9e-bce3-0b7308d2b00e · outbound

This paper cites Training language models to follow instructions with human feedback,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Training language models to follow instructions with human feedback,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.806894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.806894Z digest=sha256:aac53144e7b51f630e2626152bee328a79f3b5e6659a99ffaad605552664488f

Observation 329c671c-df29-4ec5-92ad-516a771e5422 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Direct preference optimization: Your language model is secretly a reward model,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.811299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.811299Z digest=sha256:2f376530ea160ba4cb2de0eb6a668b82c49441e3ddad40630e04b4b9dbe07ca0

Observation 3a2cc399-9573-4da9-aefa-2645b242341f · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Moshi: a speech-text foundation model for real-time dialogue

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.816937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.816937Z digest=sha256:2a9392cbc4ed3ead1af8ae923774434249317d65789ff953acd01237ffb2b13b

Observation 06b2145c-2d85-4191-bde8-1c5a2f153906 · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.821357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.821357Z digest=sha256:cd8442c2a92fcb2bb727ad230d76d5b8364cf6f5620332070883fb9e27a43bc9

Observation 36998284-953c-477a-95da-df47762a1d60 · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.825642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.825642Z digest=sha256:42d16f7bc9c225937143e1c4c1ba9f404a66144f1d74f595f1b0bc6077ac7347

Observation 8c733e4d-e0fc-41df-9404-1b8e19304a9a · outbound

This paper cites SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.829811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.829811Z digest=sha256:5c2aee8c4b133663504fdc93d95e54d2d3a1970f675faa962cec67ab4d5e5de9

Observation 19c436fa-ff2f-4d68-8e51-6a9e950764e1 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.834345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.834345Z digest=sha256:933ee17395a1cec6aaa9dadc6c061f5867f4aa99e75e083d9b8295466bf88779

Observation fedfc8f9-0c10-4bc1-80ab-eae0ffdfc593 · outbound

This paper cites LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.838344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.838344Z digest=sha256:25b493f09c6d48f1819ed6282449065bbf03ad4dc7951ea970e78d4070e03ad0

Observation e60d1685-300a-4a71-b8d3-9d17e046fbbc · outbound

This paper cites An Embarrassingly Simple Approach for LLM with Strong ASR Capacity.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM An Embarrassingly Simple Approach for LLM with Strong ASR Capacity

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.841973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.841973Z digest=sha256:ed23857330b2812b40b86c850533a27cc9ce473834fdffb234f7a41add8a44e0

Observation c4f4b24e-79ee-4853-945e-ef4672dfdf7d · outbound

This paper cites SALMONN: towards generic hearing abilities for large language models,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SALMONN: towards generic hearing abilities for large language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.641519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.845827Z digest=sha256:ccf8b60c24e3c9690350aba0467995015578e632f76b5ba6f98fd6fe1b0cbdeb

Observation 5997242c-74a7-4943-81cf-1c028e5d6d06 · outbound

This paper cites Qwen2-Audio Technical Report.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Qwen2-Audio Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.849796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.849796Z digest=sha256:1866f73f8054d4b13dab962c1ea0e9ac38ea02445182a0416824ba156f831810

Observation 181cedd0-34f3-49bf-a163-cc409242e82e · outbound

This paper cites On decoder-only architecture for speech- to-text and large language model integration,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM On decoder-only architecture for speech- to-text and large language model integration,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.626933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.854425Z digest=sha256:458eef64d0ffb4a90e0c6c08589efc5992e3a1a7f2d7620097676ebed1b2978a

Observation 782fcbb5-d293-44d4-b352-6e7db367d78f · outbound

This paper cites COSMIC: data efficient instruction-tuning for speech in-context learn- ing,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM COSMIC: data efficient instruction-tuning for speech in-context learn- ing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.610104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.858028Z digest=sha256:004499a70a40d3958e99171a78e385ee334a271f1d301e4a7bfd733db658941c

Observation e622f2e9-4fe5-42dc-bd14-35fab6d28c79 · outbound

This paper cites Prompting large language models with speech recognition abilities,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Prompting large language models with speech recognition abilities,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.590364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.862163Z digest=sha256:edb1fcf15271a9a79d393663c72174b72d46aa3af18f7dd62f12b1b2cca2848c

Observation 7df69f19-523c-44be-8936-bd2a9a07ecf9 · outbound

This paper cites SpeechVerse: A Large-scale Generalizable Audio Language Model.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.866877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.866877Z digest=sha256:ad7606d495b93cef1d88e04f8a12977d5cae240a4dba4590fb7e72e64979db27

Observation 97d59d3c-1b84-4e41-85d6-add42654e0e4 · outbound

This paper cites High fidelity neural audio compression,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM High fidelity neural audio compression,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.870796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.870796Z digest=sha256:a55d49f1c08efdec53ac0c3c6e14099f78e0445a9a0f26495278e1d36ded2cb2

Observation 623da227-e956-4251-886e-484a7e084a54 · outbound

This paper cites High- fidelity audio compression with improved rvqgan,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM High- fidelity audio compression with improved rvqgan,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.563623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.874350Z digest=sha256:45be177bb9850d470ade393eb84da68c75d4fd1d4de08471db114da30c317855

Observation 818a281e-6dc3-40a4-a1bd-5befccdef7f2 · outbound

This paper cites Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.878311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.878311Z digest=sha256:8542538d11a9c4b3de60edbc78edc9707c2dadbcf3966958884a99e5666da082

Observation 873c4ec5-451a-4960-a1e5-be8b46640b66 · outbound

This paper cites Wavllm: Towards robust and adaptive speech large language model,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Wavllm: Towards robust and adaptive speech large language model,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.553469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.882930Z digest=sha256:21148a0aaf4cf88b589e53934390d6196ae53ce3259dcab2728528ea48ede076

Observation da070468-5174-4080-9257-9d6e12c61ff6 · outbound

This paper cites Audiochatllama: Towards general-purpose speech abilities for llms,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Audiochatllama: Towards general-purpose speech abilities for llms,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.519938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.889898Z digest=sha256:e3f415addf064509d652df189a02cbadcc1f76a2b308afff7439b4833e89d32a

Observation b6e663c6-107c-432c-9304-37c183a30bc2 · outbound

This paper cites BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM BLSP: Bootstrapping Language-Speech Pre-training via Behavior Alignment of Continuation Writing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.893673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.893673Z digest=sha256:ef4e13447d9a0c62660c11f5134fb8ca99365ea2e989aa2a9a812698dc6544f9

Observation 44f31491-4a10-408e-9d3f-60af5e18d21c · outbound

This paper cites DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM DeSTA2: Developing Instruction-Following Speech Language Model Without Speech Instruction-Tuning Data

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.897763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.897763Z digest=sha256:09b59db5007c91809a6ae5e8a467c3e2f5ffb9913334fceb40ce8763f0c566ff

Observation 36749f60-17f5-410b-810f-9da282938ede · outbound

This paper cites Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Frozen Large Language Models Can Perceive Paralinguistic Aspects of Speech

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.901872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.901872Z digest=sha256:ff2b6bcf178e0ea0efe4c484fe84cb942d201abd89bc4ff4f212fd7c96e9db25

Observation 6c5b8b81-e1a0-439a-b105-2433776bea43 · outbound

This paper cites Wav2Prompt: End-to-end speech prompt learning and task-based fine-tuning for text-based LLMs,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Wav2Prompt: End-to-end speech prompt learning and task-based fine-tuning for text-based LLMs,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.506587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.905887Z digest=sha256:8560e52883a306da147bd261739c93eac85ffe83ba366dbe37b3d3581d32d7e4

Observation 92414057-f770-43da-ad24-19aa9404c3a0 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.909775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.909775Z digest=sha256:b94e1bed71659c304e9d62ef5ce52a0a6f83ee621a80b8eca7167fc25ea3cf78

Observation 36333554-2450-44f8-8090-6b818c617999 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.914477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.914477Z digest=sha256:76d2de98411b1e211c95610241e6f98ce65e82135991cfef6f33633176f1e570

Observation 65626896-ee5e-4025-a331-63ae60dc2ac7 · outbound

This paper cites Connec- tionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Connec- tionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.490200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.919090Z digest=sha256:9be4b67b7c471aeda69bc332aaf55997680fb851950d377ca43ec41dd8fd5804

Observation 71511e59-350c-4d5c-a4fd-e8c8a0a96a34 · outbound

This paper cites CASS-NAT: CTC alignment- based single step non-autoregressive transformer for speech recognition,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM CASS-NAT: CTC alignment- based single step non-autoregressive transformer for speech recognition,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.473692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.923067Z digest=sha256:a7aa47b6ee6c81069a73c566e0bc94de05cd1d61659ffe126899995b6b9b3858

Observation e5e6650f-2ee5-4b1c-ab49-d42a5d3749a0 · outbound

This paper cites Unienc-cassnat: An encoder-only non-autoregressive asr for speech ssl models,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Unienc-cassnat: An encoder-only non-autoregressive asr for speech ssl models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.461458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.927088Z digest=sha256:36a233a0a94dcb7dadd6548c4a9e3f5bed74f39f4efb4bb9d105dc34bf814172

Observation 9c2802fb-909d-4e2d-b02f-3ac59ab22d71 · outbound

This paper cites Ctc-based compression for direct speech translation,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Ctc-based compression for direct speech translation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.449574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.932204Z digest=sha256:fa48c1954d60e5c4762346b7d5de8a2d1a49e919630d53eb8a8e9d4fb3ee4b9c

Observation 982d628b-f3e8-4a9b-af47-f524d73fcd9d · outbound

This paper cites CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.435311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.935963Z digest=sha256:dac3570cacad3d84ed320fa0b92bde988887227308ab9855a6abd0bdc5418c18

Observation a2d75a61-cb1c-437a-bc2b-b126ea13604c · outbound

This paper cites SpeechT5: Unified-modal encoder-decoder pre-training for spoken language processing,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SpeechT5: Unified-modal encoder-decoder pre-training for spoken language processing,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.421186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.939645Z digest=sha256:104a7cf46ea329194557165fc85c622c3c3630e159eb11a76a3f1295677f7a93

Observation 85514468-a283-4725-a254-7017d248630d · outbound

This paper cites SpeechLM: Enhanced speech pre-training with unpaired textual data,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM SpeechLM: Enhanced speech pre-training with unpaired textual data,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.404239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.943271Z digest=sha256:d85e467784c92898325ee97dbe722570762aa52847729616a16087d26453622f

Observation 86f05e72-8f1e-4a42-a720-c8b46a45e795 · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.946785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.946785Z digest=sha256:1302050dc2dabaa39d796fc22a289b7ba95b2f01e6b3a45c1da757dd2d389158

Observation e74b7024-a6db-44f6-98ec-f26174eec041 · outbound

This paper cites M-adapter: Modality adaptation for end-to-end speech-to-text translation,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM M-adapter: Modality adaptation for end-to-end speech-to-text translation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.387882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.950774Z digest=sha256:1cc89bbaee1c09fc3f7292233e53abf21197055442cfc9b4b361aa1d356fab9a

Observation c0d740d2-39c1-4905-b8be-c793fb4494d9 · outbound

This paper cites MAESTRO: Matched speech text representa- tions through modality matching,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM MAESTRO: Matched speech text representa- tions through modality matching,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.375306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.954853Z digest=sha256:1fcfb95dcf736f5157bcc4a779298e051baa66f1cf796a9e8915c023a43f1fbd

Observation c00317ff-d02d-48b0-9d02-c6ad41888208 · outbound

This paper cites Cjst: Ctc compressor based joint speech and text training for decoder-only asr,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Cjst: Ctc compressor based joint speech and text training for decoder-only asr,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.361980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.959632Z digest=sha256:725250ddb4ec343c6eace1203e5e5eb54a38fa85d42d3df6a6a4b99c693fd63f

Observation 3b3ba2a4-aa92-49ee-829b-85887d30e4fd · outbound

This paper cites Lora: Low-rank adaptation of large language models,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Lora: Low-rank adaptation of large language models,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.346175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.969315Z digest=sha256:3aec39b79c1c6fd09dd56ab300c256a97f2f1fa364a981a8fd5a0f475bf1d677

Observation 7a096d0b-5c63-45e5-9e99-4ca0b139bb8c · outbound

This paper cites CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR

Reference 43

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T04:44:26.028455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.964450Z digest=sha256:57ba0445e6d89764a1d69d4bd3765f825e9ac945c94aca1de3df3d08fa9505b2

Observation 8c2d6d40-f604-424b-abfb-076846be80d0 · outbound

This paper cites Conformer: Convolution-augmented transformer for speech recognition,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Conformer: Convolution-augmented transformer for speech recognition,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.310290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.977083Z digest=sha256:b7e983cfdd02c67204ee8d016af04762e1f3e9059cdc899455b01b5478c6a634

Observation d6c1a48e-e5b1-4d07-9ecf-d18ee2837067 · outbound

This paper cites A CTC alignment-based non- autoregressive transformer for end-to-end automatic speech recognition,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM A CTC alignment-based non- autoregressive transformer for end-to-end automatic speech recognition,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.332237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.973181Z digest=sha256:e90ebf8931f4649ec99c72dc9127e3a4963b958fd2dab9074d659d02bcaa20ff

Observation 490ff960-f90d-44d6-a8f0-fde17f19def0 · outbound

This paper cites Unsu- pervised cross-lingual representation learning at scale,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Unsu- pervised cross-lingual representation learning at scale,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.281353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.985859Z digest=sha256:098ffbcd70191745bb15cc4f831440b3895029a2121b3b4979164bf7c3649e77

Observation 4e3c3e89-b628-42c0-b2c0-3cfcd5f1b980 · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models,.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM Zero: Memory optimizations toward training trillion parameter models,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:25.980965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:44:25.980965Z digest=sha256:e8868900c5fd1a6bede2a91a5546f7847982aa568c78791a2373d2b1fceafa35

Observation 7a11f7ff-e3ee-4275-81e7-b331e95a2782 · outbound

This paper cites 4552–4572.

AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM 4552–4572

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:44:26.537792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T04:44:25.886366Z digest=sha256:8c98ea568cfa9c57492f453088f6c1ee12314b1c6d0c3200cdc1b3743267fd6b

Pith citing papers

Observation 365b54bc-1661-477d-8383-51bc92147390 · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey AlignFormer: Modality Matching Can Achieve Better Zero-shot Instruction-Following Speech-LLM

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:45:08.151044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:fb3bfdbc8647ebdd26d8dad7fefe2b9fb5a943ff25471cd91739677d511b5d89