Pith. sign in

Paper Citation Record · LEDGER

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers

As of 9 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 1 inbound Pith citation observation for arXiv:2502.07436.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07436 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:50:14.759452Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T23:29:02.457697Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation ab335bab-23d2-4e4f-ad30-f0bae640aea8 · outbound

This paper cites GPT-4 Technical Report.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.653697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.653697Z digest=sha256:7fd885f810804120be25dfc9872e3d24f02a2e398d51e3fcc50da0783f20751d

Observation e7f6d786-4436-4c22-983f-bab6d53ca99f · outbound

This paper cites Language Models are Few-Shot Learners.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.674642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.674642Z digest=sha256:4a26c620c6bdd53a40a6a64c9649fb63a6ece39db861a4f6bd0498d7b9571e7c

Observation 9068f8ac-0ddf-4406-b958-8d7da0fa1248 · outbound

This paper cites The Llama 3 Herd of Models.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.679833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.679833Z digest=sha256:30b68e1b04421284163bc9ac54398d06449615990ed7751e0c9fc665e1727cfa

Observation f78b8bee-6c67-4fe5-ad22-3bce9a60a19f · outbound

This paper cites For language pretraining tasks, we trained LLaMA models on the BabyLM dataset.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers For language pretraining tasks, we trained LLaMA models on the BabyLM dataset

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:50:15.062540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:50:14.754909Z digest=sha256:c1879b9c94112dd462e093aa3b1661385735be8922d4dd428c3605e032ddff38

Observation 64ee1a55-faa9-42fc-9d83-6d4ad651d957 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Distilling the Knowledge in a Neural Network

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.689647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.689647Z digest=sha256:ea7e23c8223dd9521a9e5818b2d14086d5fee8c23d61b03bd4d0006798f1e840

Observation e4387d8a-7692-4398-be93-7dd99ac5b25b · outbound

This paper cites 11 Submission and Formatting Instructions for ICML 2025 Figure.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers 11 Submission and Formatting Instructions for ICML 2025 Figure

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:50:15.046119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:50:14.759452Z digest=sha256:8ee0be5b161965cd674a9e0f23254b0d8a64bbf8aacbb99b9302b48e3c9dc801

Observation c4c4921f-a0c1-45ed-be5e-db6e6fb75890 · outbound

This paper cites Improved Precision and Recall Metric for Assessing Generative Models.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Improved Precision and Recall Metric for Assessing Generative Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.704900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.704900Z digest=sha256:f9131899ebc0ef9d74ae9c46cbaf20e582ad0a31d99a103077b78b761a7f6b16

Observation 83aac50c-a978-43d1-abf3-aa3bbc9611b5 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.720671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.720671Z digest=sha256:93f37b5e7a32fd8d33d9c505002d189d7898a35d56014a052ed85712874f2514

Observation 3a7de572-72d7-4a51-80b2-fa93d0779bc4 · outbound

This paper cites Patient Knowledge Distillation for BERT Model Compression.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Patient Knowledge Distillation for BERT Model Compression

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.725290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.725290Z digest=sha256:a0423da4cc4a4354e7c2f318d652ac067b6fae7cb114a3e5aec9ae3692cfac8b

Observation a1a4deb7-9465-4e06-8cf2-3d522917c86e · outbound

This paper cites MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.730479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.730479Z digest=sha256:38bfe3f326a4c26efee6405a3184c8978b910b10b44828c2d298ce251a8a601c

Observation c86c0533-4ef1-46a9-b04b-f9bc814c9412 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.739767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.739767Z digest=sha256:3f33a90d5265cd47a98dd64c72504cd856f6cf751ac41d37d584d433941ca49d

Observation 6f1e988b-8ccc-4f1b-872a-b7e935942d91 · outbound

This paper cites Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.744931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.744931Z digest=sha256:34563fbfba4cb21cf1b63f203dc114e8237e929ca06658bc86b8ca34e8631de4

Observation ad6542bb-8d4f-49ac-9436-755170fbc72c · outbound

This paper cites ViTKD: Practical Guidelines for ViT feature knowledge distillation.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers ViTKD: Practical Guidelines for ViT feature knowledge distillation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.749967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.749967Z digest=sha256:156e30b77aa91a43eeebfc32158c7d67c529c61144c08dacf6ef16b5d037e322

Observation c1b7c2b1-06b6-4e3b-9cef-0ebb3902576c · outbound

This paper cites Like What You Like: Knowledge Distill via Neuron Selectivity Transfer.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Like What You Like: Knowledge Distill via Neuron Selectivity Transfer

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.694727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.694727Z digest=sha256:a9807a9a0dcd10935fce896cffefd9f0beeb5b693f83fe37b7c27fef77a5ea06

Observation ce1c0edc-2323-4c8f-8410-c35cf4074639 · outbound

This paper cites TinyBERT: Distilling BERT for Natural Language Understanding.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers TinyBERT: Distilling BERT for Natural Language Understanding

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.699745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.699745Z digest=sha256:7e0d28d51c6e1687960f489b83de3277e88544ba2c179410d36ee36c05933a26

Observation 0f6afcf6-867a-4516-9a2e-ebf2d6212a11 · outbound

This paper cites Analyzing and Interpreting Neural Networks for NLP: A Report on the First BlackboxNLP Workshop.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Analyzing and Interpreting Neural Networks for NLP: A Report on the First BlackboxNLP Workshop

Reference 2019

Resolution
verified exact
local_arxiv, observed 2026-08-08T12:50:14.996941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:50:14.669993Z digest=sha256:638696e54ad0b87955fa0aa32acca5bf3f85c0aac6f2543076b76978c55a8127

Observation 4f8a83dc-2af1-42c6-8018-1a4e84847cda · outbound

This paper cites Training data-efficient image transform- ers & distillation through attention.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Training data-efficient image transform- ers & distillation through attention

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:50:15.077756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:50:14.735333Z digest=sha256:74574116fd6f20c0d3c5b74fe0b7e2f7a9e819cb7339c533b97273ddb5355f9b

Observation e1812a03-25c2-4ecc-b1cf-253c2d026f34 · outbound

This paper cites Esser, P., Kulal, S., Blattmann, A., Entezari, R., M¨uller, J., Saini, H., Levi, Y ., Lorenz, D., Sauer, A., Boesel, F., et al.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers Esser, P., Kulal, S., Blattmann, A., Entezari, R., M¨uller, J., Saini, H., Levi, Y ., Lorenz, D., Sauer, A., Boesel, F., et al

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:50:15.093065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T12:50:14.684755Z digest=sha256:95a3f6be2299f180bda4e6bf17bd343209bf9f8c50ac9934391afd46791d0100

Observation 71d1e575-f4af-4a1e-8231-19ee730fd4bc · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.659381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.659381Z digest=sha256:536a445d3f77cbb73a70d296757b83d3b03d347f644f1dcf8f2cae1989f89c62

Observation 8af8114e-bc56-40e5-9dba-d72d80ca3b10 · outbound

This paper cites $V_kD:$ Improving Knowledge Distillation using Orthogonal Projections.

Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers $V_kD:$ Improving Knowledge Distillation using Orthogonal Projections

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T12:50:14.710201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:50:14.710201Z digest=sha256:57f6d8d92dcc7f4d6720eeb52e6ed62b99558aa4d76ae65038ed23b89cc4592d

Pith citing papers

Observation 66acbb21-a70f-4f7b-a051-e73e82b6bed4 · inbound

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers cites this paper.

Contribution Weights: A Geometrical Analysis of Self-Attention Transformers Optimizing Knowledge Distillation in Transformers: Enabling Multi-Head Attention without Alignment Barriers

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-06-28T23:32:46.667666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T23:29:02.457697Z digest=sha256:dfe4c8ea0026316ef21a32b2935e0b7278a57b2f163efb1630537e535fb5830b