Pith. sign in

Paper Citation Record · LEDGER

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA

As of 13 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2412.20677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20677 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:19:23.112017Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7151ea51-0361-4590-87fa-d8dec4883b12 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:22.995212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:22.995212Z digest=sha256:9a54e80a399204acc91ec1928f9501158cd7637d1cf633849646ab20e6ed6486

Observation 7fbde293-76e1-4a51-b55a-3fe3ccb88c84 · outbound

This paper cites an unresolved cited work.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:19:23.525027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:19:23.107616Z digest=sha256:df4a4af5c8a46f0520370d5154749203a82a311ad9a7cf9171c39db88af838b7

Observation 7fe3f08a-d2e1-4909-a378-f717391ae2a3 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA LoRA: Low-Rank Adaptation of Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.014442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.014442Z digest=sha256:e556c7d04f103594f9bd2864218fe5e580fa955aa125f67feb3e914e8982f3dd

Observation 343c113a-3b63-4898-b0d8-2cb8df6d0ffa · outbound

This paper cites Mistral 7B.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Mistral 7B

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.019925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.019925Z digest=sha256:64f6e65ba18c0a857667b284d4a8ad2b8c1a7039ecf5febabb7f3401a7206e55

Observation acebbc93-a900-4ccf-87e8-8f88b18c52f5 · outbound

This paper cites BiLD: Bi-directional Logits Difference Loss for Large Language Model Distillation.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA BiLD: Bi-directional Logits Difference Loss for Large Language Model Distillation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.024524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.024524Z digest=sha256:7554977bea4f20f5c847b7c8f38362cdef958737e831a17353cd3113071c86f8

Observation a489ebb9-8328-4c56-8ea7-5a33c9091de7 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.028493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.028493Z digest=sha256:00a519db7afb2ccd1b314b0480e70d840f24e0a3f6b6ff3896902244194f20ec

Observation 3b049843-b8fb-466d-a225-3f858b14dfa8 · outbound

This paper cites Learning Sparse Neural Networks through $L_0$ Regularization.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Learning Sparse Neural Networks through $L_0$ Regularization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.032946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.032946Z digest=sha256:36a75c1b759c7f4f89cf9b6ec23d62fa9255ac878ae75b5a10501d07685c1a8f

Observation b7705666-0266-47cd-b997-559697d8ccb8 · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA SocialIQA: Commonsense Reasoning about Social Interactions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.040939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.040939Z digest=sha256:21809ffc6dd7a1b75de8d68f2c6425cb99a612a5ab40544d8000af6484c28cba

Observation 3f5afc6e-3c90-4e81-97b4-c96794f439e1 · outbound

This paper cites Recursive deep models for se- mantic compositionality over a sentiment treebank.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Recursive deep models for se- mantic compositionality over a sentiment treebank

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:23.552556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:19:23.049613Z digest=sha256:14e9dbc98510bd695f352eb77cb7e29669d1782b0302021343f73cc76a137cb5

Observation 45404514-2a8d-4b1b-9869-eafcc5300f46 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA A Simple and Effective Pruning Approach for Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.053957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.053957Z digest=sha256:8d555295c01596a3a2a1a114ad0ce821b1a1e955cd8238c60ac39560f32f1978

Observation d7f1e43b-321b-4635-a518-8a98c59e813f · outbound

This paper cites RazorAttention: Efficient KV Cache Compression Through Retrieval Heads.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA RazorAttention: Efficient KV Cache Compression Through Retrieval Heads

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.058774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.058774Z digest=sha256:f9e9291eef635166c3c40aec429bb2be54216956af43141aa48879169eaee6a5

Observation b62e7200-aa4a-4488-a2bc-c7cbe7f77dd2 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.063117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.063117Z digest=sha256:dad378d111133d1e682f54c3875719bbe77f7c7b33c7e8af119e7c77767ca49a

Observation a5f7f156-3956-4295-bb7e-95440dce08e5 · outbound

This paper cites Structured Pruning of Large Language Models.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Structured Pruning of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.067137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.067137Z digest=sha256:18b027f95c40f7aef01ee6dcad3c54702256db1447e219328c77ff7ab3f8f436

Observation 62655e4c-e53d-406b-8e09-a916ae36eb06 · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.075700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.075700Z digest=sha256:49068898e4b6584697c36d595114b1726d3a764b37f15652e338e738301146dc

Observation 948d91bc-580f-4a81-9f94-579e967b72e8 · outbound

This paper cites Qwen2 Technical Report.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Qwen2 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.080312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.080312Z digest=sha256:5a7e069de65a335622de5c651ab03519a21d9fcfb99906efd323c3f55322d4f7

Observation e49414fa-1196-4d67-903c-894eef64356b · outbound

This paper cites Effectively Compress KV Heads for LLM.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Effectively Compress KV Heads for LLM

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.084595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.084595Z digest=sha256:e12e42fe142cf44f57f3dcf575f30d8f0678e34562b3519138ee4437077717d0

Observation 674a44b1-e114-4054-9095-f3990fc287f5 · outbound

This paper cites WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA WKVQuant: Quantizing Weight and Key/Value Cache for Large Language Models Gains More

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.089199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.089199Z digest=sha256:3a3d8f7277a80917a15d6868cb6467053352592eeac93420ef45b8090d94fe1a

Observation 831db064-129f-4930-bd18-0ab67374e58c · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.093364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.093364Z digest=sha256:ccdef1edf8c35832df6d0b65c3487b26ae6aa508d6b319920dba0a5948635dd4

Observation 780554e5-c090-4460-85f1-7df43cb64211 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA TinyLlama: An Open-Source Small Language Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.097981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.097981Z digest=sha256:6874d5ffef306ac77921b8784c74f34e240ce53c1292c6ab527fc3c24d1151e0

Observation 61840526-4f67-4146-bf65-83deef55c602 · outbound

This paper cites During the pruning training process, the sparsity warm-up steps account for 30% of the total steps, during which the target size of the L0 masks decreases linearly to zero.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA During the pruning training process, the sparsity warm-up steps account for 30% of the total steps, during which the target size of the L0 masks decreases linearly to zero

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:23.539292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:19:23.102598Z digest=sha256:40ee54332bc9e92a5d600d731de316552504c9bcb1087b237b4eba6e2ecf0256

Observation eac5dce5-d74b-4b33-85f6-e1199cb4fd14 · outbound

This paper cites an unresolved cited work.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:19:23.511078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:19:23.112017Z digest=sha256:7080bdc460f28ee8e7bd3ba28dd8e113164e1907cf08e8c3e8d663617bfae714

Observation ab5f951b-3fd8-45ef-87d6-b0e6d373263d · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Fast Transformer Decoding: One Write-Head is All You Need

Reference 1966

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.045321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.045321Z digest=sha256:ac8b0a53df78fef5909780fce6805129286d4a590d2f6573e50346fffd139a35

Observation b3067072-352b-47bb-8ca1-7faa285b5a07 · outbound

This paper cites DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA DHA: Learning Decoupled-Head Attention from Transformer Checkpoints via Adaptive Heads Fusion

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:22.990805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:22.990805Z digest=sha256:2a4a169d11b39db4c7686fe40c41bf47730efa88435f5ba3583ea684916a56a9

Observation 9b327448-53b0-48e8-b877-aeb244b96897 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.037009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.037009Z digest=sha256:5b399a7fec804d41f5683901399fb60f96bb9eca1cbca424401ebb813402285c

Observation d5bbf37a-8162-4973-9119-fb744a724811 · outbound

This paper cites The Llama 3 Herd of Models.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA The Llama 3 Herd of Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.004497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.004497Z digest=sha256:7de8cbdbe370a05a78924ede6120fbcd3a62fd157b0c465d83d9ef28b821c869

Observation 2f2fc995-ba2a-41ca-8213-8f1f88cd251e · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:22.999949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:22.999949Z digest=sha256:3f8e9276e1acf7089d1dd9a2f47a15f17db49902b51baa684043a4b1bd050380

Observation cf5d0739-4783-4149-ab55-ad05197b6b36 · outbound

This paper cites Lan- guage models are few-shot learners.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Lan- guage models are few-shot learners

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:19:23.566901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:19:22.986705Z digest=sha256:671eb8ac9a3a6f179c5fa31c33fdc0ea81356bf005fc7c6bdbb8c543de901adb

Observation ad30bf8b-f329-46cd-9743-077854a87204 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Measuring Massive Multitask Language Understanding

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.009235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.009235Z digest=sha256:94b1332b4df2245d4c2a6a6d32db7c8542affef81283302b1ec15ea64714158f

Observation 6aaea453-9a7a-4c1b-810d-ebbda686b8bd · outbound

This paper cites Structured Pruning Learns Compact and Accurate Models.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA Structured Pruning Learns Compact and Accurate Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:23.071189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:23.071189Z digest=sha256:f7c061637567f0a5ce8769b91a392b388768a858aefcb78b6c800f50068913e0

Observation a5f98dbc-0d0f-4192-a531-e065ae2d4a72 · outbound

This paper cites SliceGPT: Compress Large Language Models by Deleting Rows and Columns.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA SliceGPT: Compress Large Language Models by Deleting Rows and Columns

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:22.982078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:22.982078Z digest=sha256:b9ec1a07809272c8fd256896040d166fc91c1d79ccbd79838f00309e51f3e43c

Observation 2aca0bbe-26d1-4b46-a4b5-fbdb433eec82 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Align Attention Heads Before Merging Them: An Effective Way for Converting MHA to GQA GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T23:19:22.976195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:19:22.976195Z digest=sha256:8c8e1bfa44e1d926f3f0f4a5861e54f0126f0a7a5b7951427ff2e838eb805a9a

Pith citing papers

No inbound Pith citation observations are available.