Pith. sign in

Paper Citation Record · LEDGER

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers

As of 9 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 1 inbound Pith citation observation for arXiv:2502.07436.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07436 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:50:14.759452Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T23:29:02.457697Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ab335bab-23d2-4e4f-ad30-f0bae640aea8 · outbound

This paper cites GPT-4 Technical Report.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.653697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.653697Z digest=sha256:c9d722fba599b77a4b551d1b4f99769e32048ef147d634de7cfba5bdc66ed74c

Observation e7f6d786-4436-4c22-983f-bab6d53ca99f · outbound

This paper cites Language Models are Few-Shot Learners.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.674642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.674642Z digest=sha256:2ff6ba7960bb725da7c0879e4c85cf0c53f2896536ea7e717182414c58b1d3a6

Observation 9068f8ac-0ddf-4406-b958-8d7da0fa1248 · outbound

This paper cites The Llama 3 Herd of Models.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.679833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.679833Z digest=sha256:a5cf2c3147a10fc5e2c14d4b8e11cab68727f2329c2db098b339f3a885743a8d

Observation f78b8bee-6c67-4fe5-ad22-3bce9a60a19f · outbound

This paper cites For language pretraining tasks, we trained LLaMA models on the BabyLM dataset.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers For language pretraining tasks, we trained LLaMA models on the BabyLM dataset

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:50:15.062540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:50:14.754909Z digest=sha256:7307a74743fb9b6ab7164b1e6484e410e24f33adba6038fcfabf85e314eac301

Observation 64ee1a55-faa9-42fc-9d83-6d4ad651d957 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Distilling the Knowledge in a Neural Network

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.689647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.689647Z digest=sha256:320ed9d6a5d224ec82e7785ea995d068843ebb5ed5abe97885db7ee2e6dccfc7

Observation e4387d8a-7692-4398-be93-7dd99ac5b25b · outbound

This paper cites 11 Submission and Formatting Instructions for ICML 2025 Figure.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers 11 Submission and Formatting Instructions for ICML 2025 Figure

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:50:15.046119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:50:14.759452Z digest=sha256:9729dff548b17ab6e133535e4bbae03ca7461e1ad8ded76093c0e02dc20467b8

Observation c4c4921f-a0c1-45ed-be5e-db6e6fb75890 · outbound

This paper cites Improved Precision and Recall Metric for Assessing Generative Models.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Improved Precision and Recall Metric for Assessing Generative Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.704900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.704900Z digest=sha256:43ec776bf030562e4e31ea861b6817240326bca672e3281844a950f39541a26d

Observation 83aac50c-a978-43d1-abf3-aa3bbc9611b5 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.720671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.720671Z digest=sha256:997fae487aaf9b578406b2c57de79d63063e9fa14d8e06740c85108e122c7081

Observation 3a7de572-72d7-4a51-80b2-fa93d0779bc4 · outbound

This paper cites Patient Knowledge Distillation for BERT Model Compression.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Patient Knowledge Distillation for BERT Model Compression

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.725290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.725290Z digest=sha256:4e278c8835d09c323378f3b98478081b4ea41c58adbe620daa012fa5d8096721

Observation a1a4deb7-9465-4e06-8cf2-3d522917c86e · outbound

This paper cites MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.730479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.730479Z digest=sha256:03483f2933e6a25cc70d413c2d40d87fa12163c04f280b7de4d4555e6bf648d0

Observation c86c0533-4ef1-46a9-b04b-f9bc814c9412 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.739767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.739767Z digest=sha256:8e8304b1e843eb8777a3c383b70f492f54e955978ec3eea1dc573e94821c7777

Observation 6f1e988b-8ccc-4f1b-872a-b7e935942d91 · outbound

This paper cites Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.744931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.744931Z digest=sha256:1b95f765c89c2cbf03a24397515648000abddc61a3c18a9e5ef0c179dab2d541

Observation ad6542bb-8d4f-49ac-9436-755170fbc72c · outbound

This paper cites ViTKD: Practical Guidelines for ViT feature knowledge distillation.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers ViTKD: Practical Guidelines for ViT feature knowledge distillation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.749967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.749967Z digest=sha256:ab1cda103a27c74ac3e929d8a9e5c98c3208c2faa4622a199fc571551a95f502

Observation c1b7c2b1-06b6-4e3b-9cef-0ebb3902576c · outbound

This paper cites Like What You Like: Knowledge Distill via Neuron Selectivity Transfer.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Like What You Like: Knowledge Distill via Neuron Selectivity Transfer

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.694727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.694727Z digest=sha256:c5d57e5aa68c70abee832f3de139aa24facf65bad39e2a1f7bcc5eb01a647814

Observation ce1c0edc-2323-4c8f-8410-c35cf4074639 · outbound

This paper cites TinyBERT: Distilling BERT for Natural Language Understanding.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers TinyBERT: Distilling BERT for Natural Language Understanding

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.699745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.699745Z digest=sha256:eefefc3bcc822cbc7950d4c76564f0d05e93e40c840b5db3e28cec2a8e9f1678

Observation 0f6afcf6-867a-4516-9a2e-ebf2d6212a11 · outbound

This paper cites Analyzing and Interpreting Neural Networks for NLP: A Report on the First BlackboxNLP Workshop.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Analyzing and Interpreting Neural Networks for NLP: A Report on the First BlackboxNLP Workshop

Reference 2019

Resolution
verified exact
local_arxiv, observed 2026-08-08T12:50:14.996941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:50:14.669993Z digest=sha256:ea91084b9311b009b0563ab3cf5d023b3364b0697c97aec034b0d70102bad025

Observation 4f8a83dc-2af1-42c6-8018-1a4e84847cda · outbound

This paper cites Training data-efficient image transform- ers & distillation through attention.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Training data-efficient image transform- ers & distillation through attention

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:50:15.077756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:50:14.735333Z digest=sha256:63c91fb15e8becf3785cb3b9eeec49d4026fe15cc73559c0f0d2c4d88c5188c2

Observation e1812a03-25c2-4ecc-b1cf-253c2d026f34 · outbound

This paper cites Esser, P., Kulal, S., Blattmann, A., Entezari, R., M¨uller, J., Saini, H., Levi, Y ., Lorenz, D., Sauer, A., Boesel, F., et al.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Esser, P., Kulal, S., Blattmann, A., Entezari, R., M¨uller, J., Saini, H., Levi, Y ., Lorenz, D., Sauer, A., Boesel, F., et al

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:50:15.093065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:50:14.684755Z digest=sha256:de1e1aeb9583d82d976b8bd53d68d2bc1ec4c22071f8c075dd649113706dd28a

Observation 71d1e575-f4af-4a1e-8231-19ee730fd4bc · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.659381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.659381Z digest=sha256:e2520e404827b2259c80b3469f03dff64712314de71cf17f92807f57e13350f5

Observation 8af8114e-bc56-40e5-9dba-d72d80ca3b10 · outbound

This paper cites $V_kD:$ Improving Knowledge Distillation using Orthogonal Projections.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers $V_kD:$ Improving Knowledge Distillation using Orthogonal Projections

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.710201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.710201Z digest=sha256:93da68ad6063756d68e05c78063655a19419abae3486c45111ec430ea5d16c71

Pith citing papers

Observation 66acbb21-a70f-4f7b-a051-e73e82b6bed4 · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:32:46.667666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:1551fbcf5d99b9baa5b4c66274a653a726e9632d23205fea811098515fe7b137