Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:19:23.112017Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2412.20677.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:19:23.112017Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7151ea51-0361-4590-87fa-d8dec4883b12 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fbde293-76e1-4a51-b55a-3fe3ccb88c84 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7fe3f08a-d2e1-4909-a378-f717391ae2a3 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA LoRA: Low-Rank Adaptation of Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 343c113a-3b63-4898-b0d8-2cb8df6d0ffa · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Mistral 7B
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acebbc93-a900-4ccf-87e8-8f88b18c52f5 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA BiLD: Bi-directional Logits Difference Loss for Large Language Model Distillation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a489ebb9-8328-4c56-8ea7-5a33c9091de7 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b049843-b8fb-466d-a225-3f858b14dfa8 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Learning Sparse Neural Networks through $L_0$ Regularization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7705666-0266-47cd-b997-559697d8ccb8 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA SocialIQA: Commonsense Reasoning about Social Interactions
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f5afc6e-3c90-4e81-97b4-c96794f439e1 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Recursive deep models for se- mantic compositionality over a sentiment treebank
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 45404514-2a8d-4b1b-9869-eafcc5300f46 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA A Simple and Effective Pruning Approach for Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7f1e43b-321b-4635-a518-8a98c59e813f · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA RazorAttention: Efficient KV Cache Compression Through Retrieval Heads
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b62e7200-aa4a-4488-a2bc-c7cbe7f77dd2 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5f7f156-3956-4295-bb7e-95440dce08e5 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Structured Pruning of Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62655e4c-e53d-406b-8e09-a916ae36eb06 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 948d91bc-580f-4a81-9f94-579e967b72e8 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Qwen2 Technical Report
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e49414fa-1196-4d67-903c-894eef64356b · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Effectively Compress KV Heads for LLM
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 674a44b1-e114-4054-9095-f3990fc287f5 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 831db064-129f-4930-bd18-0ab67374e58c · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 780554e5-c090-4460-85f1-7df43cb64211 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA TinyLlama: An Open-Source Small Language Model
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61840526-4f67-4146-bf65-83deef55c602 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA During the pruning training process, the sparsity warm-up steps account for 30% of the total steps, during which the target size of the L0 masks decreases linearly to zero
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation eac5dce5-d74b-4b33-85f6-e1199cb4fd14 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ab5f951b-3fd8-45ef-87d6-b0e6d373263d · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Fast Transformer Decoding: One Write-Head is All You Need
Reference 1966
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3067072-352b-47bb-8ca1-7faa285b5a07 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b327448-53b0-48e8-b877-aeb244b96897 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5bbf37a-8162-4973-9119-fb744a724811 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA The Llama 3 Herd of Models
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f2fc995-ba2a-41ca-8213-8f1f88cd251e · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf5d0739-4783-4149-ab55-ad05197b6b36 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Lan- guage models are few-shot learners
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ad30bf8b-f329-46cd-9743-077854a87204 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Measuring Massive Multitask Language Understanding
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6aaea453-9a7a-4c1b-810d-ebbda686b8bd · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Structured Pruning Learns Compact and Accurate Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5f98dbc-0d0f-4192-a531-e065ae2d4a72 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA SliceGPT: Compress Large Language Models by Deleting Rows and Columns
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aca0bbe-26d1-4b46-a4b5-fbdb433eec82 · outbound
Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.