{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:G7H4UVCL5LAEO6CSDH2BF5QRCZ","short_pith_number":"pith:G7H4UVCL","schema_version":"1.0","canonical_sha256":"37cfca544beac047785219f412f6111664c2295d1f7b7e8559bcc90b436e6ae2","source":{"kind":"arxiv","id":"2402.08846","version":1},"attestation_state":"computed","paper":{"title":"An Embarrassingly Simple Approach for LLM with Strong ASR Capacity","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.MM","cs.SD","eess.AS"],"primary_cat":"cs.CL","authors_text":"Fan Yu, Guanrou Yang, Jiaming Wang, Qian Chen, Shiliang Zhang, Siqi Zheng, Xie Chen, Yifan Yang, Zhifu Gao, Zhihao Du, Ziyang Ma","submitted_at":"2024-02-13T23:25:04Z","abstract_excerpt":"In this paper, we focus on solving one of the most important tasks in the field of speech processing, i.e., automatic speech recognition (ASR), with speech foundation encoders and large language models (LLM). Recent works have complex designs such as compressing the output temporally for the speech encoder, tackling modal alignment for the projector, and utilizing parameter-efficient fine-tuning for the LLM. We found that delicate designs are not necessary, while an embarrassingly simple composition of off-the-shelf speech encoder, LLM, and the only trainable linear projector is competent for "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2402.08846","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2024-02-13T23:25:04Z","cross_cats_sorted":["cs.AI","cs.MM","cs.SD","eess.AS"],"title_canon_sha256":"2616b9ecfdc3de8cd7b6a0595d2d2fea2021dacf16f1deced7723081a30ea7b5","abstract_canon_sha256":"b23ae9eac08a3ed7114141c489a353cfbd9624ca7c570cd5ee518991c9b81e8f"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T07:45:03.586697Z","signature_b64":"4cq0gowQ2Yx2ICkANlfJFiAGL6CPaDzacpg4/wI5sZ0SaaEDWdg977+vpn+Xgky0//quJ8ilAQ9LpnJFPIqyAA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"37cfca544beac047785219f412f6111664c2295d1f7b7e8559bcc90b436e6ae2","last_reissued_at":"2026-07-05T07:45:03.586198Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T07:45:03.586198Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"An Embarrassingly Simple Approach for LLM with Strong ASR Capacity","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"","cross_cats":["cs.AI","cs.MM","cs.SD","eess.AS"],"primary_cat":"cs.CL","authors_text":"Fan Yu, Guanrou Yang, Jiaming Wang, Qian Chen, Shiliang Zhang, Siqi Zheng, Xie Chen, Yifan Yang, Zhifu Gao, Zhihao Du, Ziyang Ma","submitted_at":"2024-02-13T23:25:04Z","abstract_excerpt":"In this paper, we focus on solving one of the most important tasks in the field of speech processing, i.e., automatic speech recognition (ASR), with speech foundation encoders and large language models (LLM). Recent works have complex designs such as compressing the output temporally for the speech encoder, tackling modal alignment for the projector, and utilizing parameter-efficient fine-tuning for the LLM. We found that delicate designs are not necessary, while an embarrassingly simple composition of off-the-shelf speech encoder, LLM, and the only trainable linear projector is competent for "},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2402.08846","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2402.08846/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2402.08846","created_at":"2026-07-05T07:45:03.586256+00:00"},{"alias_kind":"arxiv_version","alias_value":"2402.08846v1","created_at":"2026-07-05T07:45:03.586256+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2402.08846","created_at":"2026-07-05T07:45:03.586256+00:00"},{"alias_kind":"pith_short_12","alias_value":"G7H4UVCL5LAE","created_at":"2026-07-05T07:45:03.586256+00:00"},{"alias_kind":"pith_short_16","alias_value":"G7H4UVCL5LAEO6CS","created_at":"2026-07-05T07:45:03.586256+00:00"},{"alias_kind":"pith_short_8","alias_value":"G7H4UVCL","created_at":"2026-07-05T07:45:03.586256+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":20,"internal_anchor_count":2,"sample":[{"citing_arxiv_id":"2607.06827","citing_title":"Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs","ref_index":15,"is_internal_anchor":true},{"citing_arxiv_id":"2607.08409","citing_title":"When Synthetic Speech Is All You Have: Better Call GRPO","ref_index":14,"is_internal_anchor":true},{"citing_arxiv_id":"2606.25444","citing_title":"Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs?","ref_index":31,"is_internal_anchor":false},{"citing_arxiv_id":"2606.24123","citing_title":"Aligning MusicLLM with Emotion using Instruction Tuning and Feedback-Driven Alignment","ref_index":7,"is_internal_anchor":false},{"citing_arxiv_id":"2606.10439","citing_title":"Enhancing Multilingual LLM-based ASR with Mixture of Experts and Dynamic Downsampling","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2606.10368","citing_title":"Speech Meets ELF: Audio Conditional Continuous-Target Diffusion for Speech Recognition and Translation","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2606.10454","citing_title":"Entropy-Aware Domain-Routed Mixture-of-Experts Speech-LLM Framework: A Case Study of Multi-Domain Child-Adult ASR","ref_index":10,"is_internal_anchor":false},{"citing_arxiv_id":"2606.09966","citing_title":"RespiraMFM: A Multimodal Foundation Model with Contrastive Audio-Language Alignment for Respiratory Disease Identification","ref_index":44,"is_internal_anchor":false},{"citing_arxiv_id":"2606.08486","citing_title":"TRADE: Transducer-Augmented Decoder for Speech LLM","ref_index":24,"is_internal_anchor":false},{"citing_arxiv_id":"2606.02400","citing_title":"SoulX-Transcriber: A Robust End-to-End Framework for Multi-Speaker Speech Transcription","ref_index":4,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14340","citing_title":"Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2601.20898","citing_title":"Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2603.15045","citing_title":"LLMs and Speech: Integration vs. Combination","ref_index":32,"is_internal_anchor":false},{"citing_arxiv_id":"2605.14340","citing_title":"Refining Pseudo-Audio Prompts with Speech-Text Alignment for Text-Only Domain Adaptation in LLM-Based ASR","ref_index":11,"is_internal_anchor":false},{"citing_arxiv_id":"2604.11269","citing_title":"Speaker Attributed Automatic Speech Recognition Using Speech Aware LLMS","ref_index":13,"is_internal_anchor":false},{"citing_arxiv_id":"2604.09332","citing_title":"Phonemes vs. Projectors: An Investigation of Speech-Language Interfaces for LLM-based ASR","ref_index":20,"is_internal_anchor":false},{"citing_arxiv_id":"2604.09121","citing_title":"Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition","ref_index":18,"is_internal_anchor":false},{"citing_arxiv_id":"2604.06487","citing_title":"Closing the Speech-Text Gap with Limited Audio for Effective Domain Adaptation in LLM-Based ASR","ref_index":8,"is_internal_anchor":false},{"citing_arxiv_id":"2604.12398","citing_title":"Contextual Biasing for ASR in Speech LLM with Common Word Cues and Bias Word Position Prediction","ref_index":32,"is_internal_anchor":false},{"citing_arxiv_id":"2604.22817","citing_title":"In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions","ref_index":14,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/G7H4UVCL5LAEO6CSDH2BF5QRCZ","json":"https://pith.science/pith/G7H4UVCL5LAEO6CSDH2BF5QRCZ.json","graph_json":"https://pith.science/api/pith-number/G7H4UVCL5LAEO6CSDH2BF5QRCZ/graph.json","events_json":"https://pith.science/api/pith-number/G7H4UVCL5LAEO6CSDH2BF5QRCZ/events.json","paper":"https://pith.science/paper/G7H4UVCL"},"agent_actions":{"view_html":"https://pith.science/pith/G7H4UVCL5LAEO6CSDH2BF5QRCZ","download_json":"https://pith.science/pith/G7H4UVCL5LAEO6CSDH2BF5QRCZ.json","view_paper":"https://pith.science/paper/G7H4UVCL","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2402.08846&json=true","fetch_graph":"https://pith.science/api/pith-number/G7H4UVCL5LAEO6CSDH2BF5QRCZ/graph.json","fetch_events":"https://pith.science/api/pith-number/G7H4UVCL5LAEO6CSDH2BF5QRCZ/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/G7H4UVCL5LAEO6CSDH2BF5QRCZ/action/timestamp_anchor","attest_storage":"https://pith.science/pith/G7H4UVCL5LAEO6CSDH2BF5QRCZ/action/storage_attestation","attest_author":"https://pith.science/pith/G7H4UVCL5LAEO6CSDH2BF5QRCZ/action/author_attestation","sign_citation":"https://pith.science/pith/G7H4UVCL5LAEO6CSDH2BF5QRCZ/action/citation_signature","submit_replication":"https://pith.science/pith/G7H4UVCL5LAEO6CSDH2BF5QRCZ/action/replication_record"}},"created_at":"2026-07-05T07:45:03.586256+00:00","updated_at":"2026-07-05T07:45:03.586256+00:00"}